The RunPod alternative that bills tokens, not GPU-hours
On RunPod you rent the GPU, load the model, tune the worker and pay while it boots or idles. Featherless serves 40,000+ open models on one OpenAI-compatible API and bills the tokens you use. Every RunPod figure on this page comes from RunPod's own pricing page and docs.
Four places the meter runs while you wait
RunPod is a good GPU cloud, and for training, fine-tuning, ComfyUI or anything that needs root on a machine we'd point you there ourselves. Serving an LLM behind an API is where a GPU-hour meter and a request pattern pull in different directions.
Loading the model
RunPod bills from 'initializing the container and loading models into GPU memory'. Its own engineers measured a 324-second cold start for a 32B model, and 91 seconds after four configuration changes.
Sitting idle
Flex workers bill through the idle timeout after every request, five seconds by default. Active workers bill 24/7 whether or not a request arrives, and they are the only way to avoid the cold start.
Running the stack
The vLLM worker documents more than 60 environment variables. Tool-call parsers, tensor parallelism, memory utilisation and the model download are yours to configure, and yours to fix when they break.
Finding a GPU
Stop a pod and RunPod may have rented its GPU to someone else by the time you restart. Its Deploy When Available feature exists because, in RunPod's words, 'the industry-wide GPU supply crunch has introduced an enormous amount of friction'.
RunPod pricing vs ours, model by model
RunPod sells a handful of language models by the token through Public Endpoints; every other model means a Serverless GPU worker you configure and pay for by the second. Rates as published by RunPod on 16 September 2026.
| Model | RunPod in / out | Featherless in / out | Difference |
|---|---|---|---|
| Kimi K32780B MoE · 1M ctx on RunPod, 256K here | $3.00 / $15.00Public Endpoint | $3.00 / $15.00 | level |
| Kimi K2.61T MoE · 256K ctx | $0.95 / $4.00Public Endpoint | $0.80 / $3.40 | −15% |
| Kimi K2.7 Code256K ctx | $0.95 / $4.00Public Endpoint | $0.80 / $3.40 | −15% |
| Qwen3 32BAWQ 4-bit on RunPod | $10.00 flatPublic Endpoint · per 1M tokens | $0.102 / $0.493 | −97% |
| gpt-oss-120B128K ctx | $10.00 flatPublic Endpoint · per 1M tokens | $0.15 / $0.60 | −96% |
| GLM-5.3753B · 256K ctx | Not sold by the tokenneeds your own vLLM worker | $1.40 / $4.40 | — |
| DeepSeek V4 Pro256K ctx | Not sold by the tokenneeds your own vLLM worker | $1.60 / $3.20 | — |
| DeepSeek V4 Flash284B · 256K ctx | Not sold by the tokenneeds your own vLLM worker | $0.14 / $0.28 | — |
| Qwen3.5 397B A17B | Not sold by the tokenneeds your own vLLM worker | $0.55 / $3.50 | — |
| MiniMax M3427B · 256K ctx | Not sold by the tokenneeds your own vLLM worker | $0.55 / $2.20 | — |
| Gemma 4 31B | Not sold by the tokenneeds your own vLLM worker | $0.12 / $0.36 | — |
| Llama 3.3 70B | Not sold by the tokenneeds your own vLLM worker | $0.65 / $0.75 | — |
What always-on costs on a GPU meter
GLM-5.3 behind an API, 100M input and 100M output tokens a month. RunPod doesn't sell it by the token, and the only way to avoid a multi-minute cold start after every quiet spell is an active worker, so that is what we price.
4 × B300 280GB at $9.98 per GPU-hour, 730 hours. GLM-5.3 is 753B parameters, about 750GB of weights at 8-bit before the KV cache, so the smallest worker that holds it is four B300s, or eight H200s at $5.93 for $34,631. Active workers bill continuously, idle included; the discount on them is by sales enquiry.
100M input at $1.40 = $140. 100M output at $4.40 = $440. Nothing to size, load or keep warm; the model is in the catalogue at its Hugging Face repo path.
Predictable pricing that scales with you
- Context size up to 256K
- 1 agent environment included
- Fastest response times
- Unused credits roll over
- Billed per token
- Dedicated H100, MI325, B200 & B300 GPUs
- Engineering team included
- Gets cheaper over time with fine-tuning
- Burst & failover to Public Cloud
A managed node for about a third of a bare one
Above a certain sustained volume, reserved capacity beats metered pricing on any platform. RunPod publishes Secure Cloud pods by the GPU-hour, so the comparison is arithmetic, with one difference: a RunPod pod is bare hardware, and the serving stack, quantisation and on-call are yours.
8 × H100 SXM at $3.49 per GPU-hour on Secure Cloud, running continuously (730 hrs). Community Cloud has the same node for $15,710 on peer-to-peer hosts RunPod rates 'Variable' for reliability. 8 × H200 is $26,806 and 8 × B200 is $39,654. Storage for the weights is extra.
A dedicated AMD MI325X node, reserved and isolated. Inference stack tuning, quantisation, batching and on-call support are included rather than billed separately.
3- or 6-month non-refundable term, but storage is billed at full rate throughout, and a stopped pod may not get its GPU back when you restart it.What the hardware sustains
| GPU | Memory | GPT-OSS 120B | Gemma 4 31B |
|---|---|---|---|
| AMD MI325X | 256 GB HBM3 | 500M tok/mo | 675M tok/mo |
| NVIDIA H100 | 80 GB HBM2e | 200M tok/mo | 70M tok/mo |
| NVIDIA B200 | 180 GB HBM3e | 1B tok/mo | 700M tok/mo |
| NVIDIA B300 | 288 GB HBM3e | 1.5B tok/mo | 1B tok/mo |
Switching from RunPod: a one-line change
RunPod's vLLM worker and its Public Endpoints both speak OpenAI Chat Completions, and so do we. If you're calling api.runpod.ai through the OpenAI SDK today, swap the base URL and key. The model name stays the Hugging Face repo path you already use.
Get your API keyFrequently asked questions
It depends which half of RunPod you use. For serving open LLMs behind an API, Featherless is the closest like-for-like: an OpenAI-compatible endpoint with 40,000+ models billed per token, against RunPod's five per-token language models and a Serverless tier where you bring the worker. On Kimi K3 the list rate is identical; on Kimi K2.6 and K2.7 Code we're 15% lower. For renting GPUs by the hour, training, fine-tuning, or image and video work, we're not an alternative at all and RunPod is the right tool.
No. Chat is sold for interactive, human-driven use by the subscriber, not for API traffic, background automation, reselling or benchmarking. If an application is making the calls, you want the Developer plan. We're explicit about this because signing up for the wrong plan is the single most common source of frustration when people switch.
No. Against Public Endpoints we're level on Kimi K3 and about 15% lower on Kimi K2.6 and K2.7 Code, a modest gap either way; the flat $10 per million tokens on Qwen3 32B and gpt-oss-120B is where the gap is large. Against a Serverless worker you run yourself, per-token wins at low and medium volume because the meter runs through idle time and boot time. A GPU you keep saturated around the clock can win at very high volume, and at that point a reserved node, bare or managed, is the right comparison.
For inference serving the gap with Nvidia is much narrower than it is for training, and on memory-bound workloads the MI325X's 256GB is an advantage: it sustains roughly 2.5x the GPT-OSS 120B throughput of an H100 on our infrastructure. RunPod's catalogue is almost entirely Nvidia; its one AMD part, the MI300X, is priced on its own page but not on the main pricing table. Where AMD matters is model compatibility: most of the catalogue runs on it, not every model does. Ask before you migrate a specific model and we'll tell you straight. We also run Nvidia H100, B200 and B300 if that is the better fit.
Reserved, isolated GPU capacity plus the team to operate it, inference stack tuning, quantisation, batching, fine-tuning on your traffic, and on-call support. Not a separate professional services engagement. We benchmark your workload before you commit, then guarantee that performance level on the reserved capacity.
Yes, on qualifying models, applied automatically when a request reuses a prompt prefix of roughly 1,000+ tokens we served recently: Kimi K3 $0.30, Kimi K2.6 $0.154, GLM 5.2 $0.15 and gpt-oss-120B $0.02 per million cached input tokens. RunPod's Public Endpoints publish no cached-input rate, and on a Serverless worker prefix caching is a vLLM flag you enable and pay GPU time for either way.
Multiply your monthly tokens by our per-token rate. Then price the RunPod side the way its meter will: the GPU-hours you actually hold, boot time on every scale-from-zero, the idle timeout after every request, network storage for the weights at $0.07 per GB a month, and the engineering time to build and keep the worker. Most comparisons stop at the hourly rate and skip the last four. The cost calculator will do it for you, or talk to an engineer and we'll model your real traffic.
No. We sell tokens and managed dedicated nodes, not raw GPU-hours: no pods, no per-second workers, no training clusters. Fine-tuning is available on the Business tier, on your own dedicated capacity. If root access, a training run, ComfyUI or Whisper is your main requirement, RunPod is the better fit and we'd rather say so now than after you've migrated.
Run the numbers on your own workload
Price your traffic against both, idle and boot time included, or take an API key and test on real load.