The RunPod alternative that bills tokens, not GPU-hours

On RunPod you rent the GPU, load the model, tune the worker and pay while it boots or idles. Featherless serves 40,000+ open models on one OpenAI-compatible API and bills the tokens you use. Every RunPod figure on this page comes from RunPod's own pricing page and docs.

Four places the meter runs while you wait

RunPod is a good GPU cloud, and for training, fine-tuning, ComfyUI or anything that needs root on a machine we'd point you there ourselves. Serving an LLM behind an API is where a GPU-hour meter and a request pattern pull in different directions.

◷

Loading the model

RunPod bills from 'initializing the container and loading models into GPU memory'. Its own engineers measured a 324-second cold start for a 32B model, and 91 seconds after four configuration changes.

⊞

Sitting idle

Flex workers bill through the idle timeout after every request, five seconds by default. Active workers bill 24/7 whether or not a request arrives, and they are the only way to avoid the cold start.

↕

Running the stack

The vLLM worker documents more than 60 environment variables. Tool-call parsers, tensor parallelism, memory utilisation and the model download are yours to configure, and yours to fix when they break.

◈

Finding a GPU

Stop a pod and RunPod may have rented its GPU to someone else by the time you restart. Its Deploy When Available feature exists because, in RunPod's words, 'the industry-wide GPU supply crunch has introduced an enormous amount of friction'.

RunPod pricing vs ours, model by model

RunPod sells a handful of language models by the token through Public Endpoints; every other model means a Serverless GPU worker you configure and pay for by the second. Rates as published by RunPod on 16 September 2026.

ModelRunPod
in / out
Featherless
in / out
Difference
Kimi K32780B MoE · 1M ctx on RunPod, 256K here$3.00 / $15.00Public Endpoint$3.00 / $15.00level
Kimi K2.61T MoE · 256K ctx$0.95 / $4.00Public Endpoint$0.80 / $3.40−15%
Kimi K2.7 Code256K ctx$0.95 / $4.00Public Endpoint$0.80 / $3.40−15%
Qwen3 32BAWQ 4-bit on RunPod$10.00 flatPublic Endpoint · per 1M tokens$0.102 / $0.493−97%
gpt-oss-120B128K ctx$10.00 flatPublic Endpoint · per 1M tokens$0.15 / $0.60−96%
GLM-5.3753B · 256K ctxNot sold by the tokenneeds your own vLLM worker$1.40 / $4.40—
DeepSeek V4 Pro256K ctxNot sold by the tokenneeds your own vLLM worker$1.60 / $3.20—
DeepSeek V4 Flash284B · 256K ctxNot sold by the tokenneeds your own vLLM worker$0.14 / $0.28—
Qwen3.5 397B A17BNot sold by the tokenneeds your own vLLM worker$0.55 / $3.50—
MiniMax M3427B · 256K ctxNot sold by the tokenneeds your own vLLM worker$0.55 / $2.20—
Gemma 4 31BNot sold by the tokenneeds your own vLLM worker$0.12 / $0.36—
Llama 3.3 70BNot sold by the tokenneeds your own vLLM worker$0.65 / $0.75—
Five models by the token, everything else by the hour. On Kimi K3 the list rate is identical and on Kimi K2.6 and K2.7 Code we are 15% lower, so if one of those is your model, price is not the reason to move. The flat $10 per million on Qwen3 32B and gpt-oss-120B is. So is the sixth model: anything not on Public Endpoints means a Serverless worker you size, configure and pay for by the second, boot time included. We run on AMD MI-series accelerators and load models on demand from Hugging Face weights, which is why the catalogue is 40,000 deep and none of it needs a worker of yours.

What always-on costs on a GPU meter

GLM-5.3 behind an API, 100M input and 100M output tokens a month. RunPod doesn't sell it by the token, and the only way to avoid a multi-minute cold start after every quiet spell is an active worker, so that is what we price.

RunPod Serverless — active workers
$29,142/mo

4 × B300 280GB at $9.98 per GPU-hour, 730 hours. GLM-5.3 is 753B parameters, about 750GB of weights at 8-bit before the KV cache, so the smallest worker that holds it is four B300s, or eight H200s at $5.93 for $34,631. Active workers bill continuously, idle included; the discount on them is by sales enquiry.

Token cost on Featherless for GLM-5.3
$580/mo

100M input at $1.40 = $140. 100M output at $4.40 = $440. Nothing to size, load or keep warm; the model is in the catalogue at its Hugging Face repo path.

$28,562 a month, before about $50 of network storage for the weights and the time spent tuning the worker. Flex workers cut the bill when traffic is bursty and hand you the boot instead: RunPod's own benchmark put a 32B cold start at 324 seconds, billed. At very high sustained volume a GPU you keep saturated beats any token meter, which is what the dedicated node below is for. Check your own model and volume rather than trusting either headline. Use the calculator →

Predictable pricing that scales with you

developer
Build production Al with the fastest usage-based inference.
$50 credits/month
  • Context size up to 256K
  • 1 agent environment included
  • Fastest response times
  • Unused credits roll over
  • Billed per token
Credit Amount
Subscribe
Billed monthly. Cancel anytime.
business
Dedicated GPUs and the team to run them.
Custom
  • Dedicated H100, MI325, B200 & B300 GPUs
  • Engineering team included
  • Gets cheaper over time with fine-tuning
  • Burst & failover to Public Cloud
Talk to an Engineer
Annual contracts. Volume pricing.

A managed node for about a third of a bare one

Above a certain sustained volume, reserved capacity beats metered pricing on any platform. RunPod publishes Secure Cloud pods by the GPU-hour, so the comparison is arithmetic, with one difference: a RunPod pod is bare hardware, and the serving stack, quantisation and on-call are yours.

RunPod — 8 × H100, bare pod
$20,382/mo

8 × H100 SXM at $3.49 per GPU-hour on Secure Cloud, running continuously (730 hrs). Community Cloud has the same node for $15,710 on peer-to-peer hosts RunPod rates 'Variable' for reliability. 8 × H200 is $26,806 and 8 × B200 is $39,654. Storage for the weights is extra.

Featherless — full managed node
$7,500/mo

A dedicated AMD MI325X node, reserved and isolated. Inference stack tuning, quantisation, batching and on-call support are included rather than billed separately.

Roughly 2.7x less than the bare Secure Cloud node, and 2.1x less than Community Cloud, with the operations included rather than yours to build. RunPod's savings plans bring the compute rate down on a 3- or 6-month non-refundable term, but storage is billed at full rate throughout, and a stopped pod may not get its GPU back when you restart it.

What the hardware sustains

GPUMemoryGPT-OSS 120BGemma 4 31B
AMD MI325X256 GB HBM3500M tok/mo675M tok/mo
NVIDIA H10080 GB HBM2e200M tok/mo70M tok/mo
NVIDIA B200180 GB HBM3e1B tok/mo700M tok/mo
NVIDIA B300288 GB HBM3e1.5B tok/mo1B tok/mo
How it works in practice: Available in the US, EU and Southeast Asia with no procurement cycle.

Switching from RunPod: a one-line change

RunPod's vLLM worker and its Public Endpoints both speak OpenAI Chat Completions, and so do we. If you're calling api.runpod.ai through the OpenAI SDK today, swap the base URL and key. The model name stays the Hugging Face repo path you already use.

Get your API key
# Before — RunPod Serverless (vLLM worker) client = OpenAI( base_url="https://api.runpod.ai/v2/ENDPOINT_ID/openai/v1", api_key=os.environ["RUNPOD_API_KEY"] )   # After — Featherless client = OpenAI( base_url="https://api.featherless.ai/v1", api_key=os.environ["FEATHERLESS_API_KEY"] )   # Same Hugging Face repo path, no worker to configure r = client.chat.completions.create( model="meta-llama/Llama-3.3-70B-Instruct", messages=[{"role": "user", "content": "Hello"}] )

Frequently asked questions

What is the best RunPod alternative?

It depends which half of RunPod you use. For serving open LLMs behind an API, Featherless is the closest like-for-like: an OpenAI-compatible endpoint with 40,000+ models billed per token, against RunPod's five per-token language models and a Serverless tier where you bring the worker. On Kimi K3 the list rate is identical; on Kimi K2.6 and K2.7 Code we're 15% lower. For renting GPUs by the hour, training, fine-tuning, or image and video work, we're not an alternative at all and RunPod is the right tool.

Can I use the $25 Chat plan for my application?

No. Chat is sold for interactive, human-driven use by the subscriber, not for API traffic, background automation, reselling or benchmarking. If an application is making the calls, you want the Developer plan. We're explicit about this because signing up for the wrong plan is the single most common source of frustration when people switch.

Is Featherless always cheaper than RunPod?

No. Against Public Endpoints we're level on Kimi K3 and about 15% lower on Kimi K2.6 and K2.7 Code, a modest gap either way; the flat $10 per million tokens on Qwen3 32B and gpt-oss-120B is where the gap is large. Against a Serverless worker you run yourself, per-token wins at low and medium volume because the meter runs through idle time and boot time. A GPU you keep saturated around the clock can win at very high volume, and at that point a reserved node, bare or managed, is the right comparison.

Is AMD hardware slower for inference?

For inference serving the gap with Nvidia is much narrower than it is for training, and on memory-bound workloads the MI325X's 256GB is an advantage: it sustains roughly 2.5x the GPT-OSS 120B throughput of an H100 on our infrastructure. RunPod's catalogue is almost entirely Nvidia; its one AMD part, the MI300X, is priced on its own page but not on the main pricing table. Where AMD matters is model compatibility: most of the catalogue runs on it, not every model does. Ask before you migrate a specific model and we'll tell you straight. We also run Nvidia H100, B200 and B300 if that is the better fit.

What does the $7,500/month dedicated node include?

Reserved, isolated GPU capacity plus the team to operate it, inference stack tuning, quantisation, batching, fine-tuning on your traffic, and on-call support. Not a separate professional services engagement. We benchmark your workload before you commit, then guarantee that performance level on the reserved capacity.

Do you offer cached input pricing?

Yes, on qualifying models, applied automatically when a request reuses a prompt prefix of roughly 1,000+ tokens we served recently: Kimi K3 $0.30, Kimi K2.6 $0.154, GLM 5.2 $0.15 and gpt-oss-120B $0.02 per million cached input tokens. RunPod's Public Endpoints publish no cached-input rate, and on a Serverless worker prefix caching is a vLLM flag you enable and pay GPU time for either way.

How do I work out which is cheaper for me?

Multiply your monthly tokens by our per-token rate. Then price the RunPod side the way its meter will: the GPU-hours you actually hold, boot time on every scale-from-zero, the idle timeout after every request, network storage for the weights at $0.07 per GB a month, and the engineering time to build and keep the worker. Most comparisons stop at the hourly rate and skip the last four. The cost calculator will do it for you, or talk to an engineer and we'll model your real traffic.

Do you rent GPUs by the hour like RunPod?

No. We sell tokens and managed dedicated nodes, not raw GPU-hours: no pods, no per-second workers, no training clusters. Fine-tuning is available on the Business tier, on your own dedicated capacity. If root access, a training run, ComfyUI or Whisper is your main requirement, RunPod is the better fit and we'd rather say so now than after you've migrated.

Run the numbers on your own workload

Price your traffic against both, idle and boot time included, or take an API key and test on real load.