The Fireworks AI alternative that doesn't retire your model
Fireworks removes serverless models in waves: Llama 3.3 70B in May, gpt-oss-20B in August, GLM 5.2 and Kimi K2.6 on 25 September 2026. All four are still served here, next to 40,000+ other open models, on one OpenAI-compatible API at one rate.
A serverless tier that keeps getting smaller
Fireworks is fast and well engineered, and the right choice if you want managed fine-tuning or Blackwell GPUs by the hour.
Models leave on two weeks' notice
Multiple rounds of serverless removals. The 25 September batch was announced on 12 September and takes GLM 5.2, Kimi K2.6, Kimi K2.7 Code and both DeepSeek V4 builds.
Three prices for one model
Standard is shed first under load. Priority costs 20 to 58% more, Fast costs 50% more. On Kimi K3 that is $15.00, $18.75 or $22.50 per million output tokens for the same weights.
Your fine-tune needs its own GPU
LoRAs and uploaded models can't run on Fireworks serverless. They go on a dedicated deployment from $8 per GPU-hour, and a deployment scaled to zero answers 503 while it wakes.
GPU prices rose in September
On 1 September 2026 Fireworks raised on-demand rates: H100 from $7 to $8 an hour, B200 from $10 to $13. Region-restricted deployments add another 50%.
Fireworks AI pricing vs ours, model by model
Fireworks' Standard serverless rates against ours, checked 16 September 2026. Cached input shown where Fireworks publishes it. Priority adds 20 to 58% and Fast adds 50% to every figure in the Fireworks column.
| Model | Fireworks in / out | Featherless in / out | Difference |
|---|---|---|---|
| Kimi K32780B MoE · 256K ctx | $3.00 / $15.00$0.30 cached in | $3.00 / $15.00 | level |
| GLM-5.3753B · 256K ctx | $1.40 / $4.40$0.26 cached in | $1.40 / $4.40 | level |
| GLM-5.2753B · off serverless 25 Sep 2026 | $1.40 / $4.40$0.14 cached in | $1.40 / $4.40 | level |
| Kimi K2.61T MoE · off serverless 25 Sep 2026 | $0.95 / $4.00$0.16 cached in | $0.80 / $3.40 | −15% |
| DeepSeek V4 Pro0813 build · off serverless 25 Sep 2026 | $1.32 / $3.96$0.044 cached in | $1.60 / $3.20 | −9% |
| DeepSeek V4 Flash284B · 256K ctx | $0.22 / $0.66$0.007 cached in | $0.14 / $0.28 | −52% |
| MiniMax M3427B · 256K ctx | $0.30 / $1.20$0.06 cached in | $0.55 / $2.20 | +83% |
| gpt-oss-120B128K ctx | $0.15 / $0.60$0.015 cached in | $0.15 / $0.60 | level |
| Qwen3.5 397B A17B | Not on serverlesson-demand GPU only | $0.55 / $3.50 | — |
| Gemma 4 31B | Not on serverlesson-demand GPU only | $0.12 / $0.36 | — |
| gpt-oss-20B128K ctx | Removed from serverless27 Aug 2026 | $0.04 / $0.15 | — |
| Llama 3.3 70B | Removed from serverless14 May 2026 | $0.65 / $0.75 | — |
What a deprecation costs
Fireworks removes Kimi K2.6 from serverless on 25 September 2026 and suggests GLM 5.3 or Kimi K3 instead. Say you want the model you shipped with. 100M input and 100M output tokens a month, list rates on both sides.
4 × B300 at $15.00 per GPU-hour for 730 hours. A 1T-parameter model at 8-bit needs about a terabyte of GPU memory, so the smallest deployment that holds Kimi K2.6 is four B300s, or eight H200s at $8.00 for $46,720. Scale-to-zero trims the bill but answers 503 while the replica warms, and a deployment with no traffic for 7 days is deleted.
100M input at $0.80 = $80. 100M output at $3.40 = $340. The model never left the catalogue, so nothing in your integration changes.
Predictable pricing that scales with you
- Context size up to 256K
- 1 agent environment included
- Fastest response times
- Unused credits roll over
- Billed per token
- Dedicated H100, MI325, B200 & B300 GPUs
- Engineering team included
- Gets cheaper over time with fine-tuning
- Burst & failover to Public Cloud
A managed node for a third of a Fireworks half node
Above a certain sustained volume, reserved capacity beats metered pricing on any platform. Fireworks publishes on-demand deployments at a per-GPU hourly rate, raised on 1 September 2026, so the comparison is arithmetic.
4 × H100 at $8.00 per GPU per hour, running continuously (730 hrs). A full 8-GPU node is $46,720. On B200 at $13.00/hr the same half node is $37,960, and a region-restricted deployment adds 50% on top.
A dedicated AMD MI325X node, reserved and isolated. Inference stack tuning, quantisation, batching and on-call support are included rather than billed separately.
Switching from Fireworks AI: a one-line change
Both APIs are OpenAI-compatible. If you're calling Fireworks through the OpenAI SDK today, swap the base URL and key, then replace the accounts/fireworks/models/ prefix with the Hugging Face repo path.
Get your API keyFrequently asked questions
For serverless open-model inference, Featherless is the closest like-for-like: the same OpenAI-compatible API, the same flagship models at level or lower list rates (Kimi K3 $3.00 / $15.00 on both; DeepSeek V4 Flash $0.42 against $0.88 per 1M in + 1M out, September 2026), plus 40,000+ open models that Fireworks serves only on rented GPUs or not at all. Fireworks remains the better choice for managed fine-tuning, batch inference at half price, speculative decoding on dedicated deployments, and Blackwell GPUs by the hour.
No. Chat is sold for interactive, human-driven use by the subscriber, not for API traffic, background automation, reselling or benchmarking. If an application is making the calls, you want the Developer plan. We're explicit about this because signing up for the wrong plan is the single most common source of frustration when people switch.
On list rates, mostly level. We match Fireworks Standard on Kimi K3, GLM 5.2, GLM 5.3 and gpt-oss-120B, undercut it on Kimi K2.6 (15% lower) and DeepSeek V4 Flash (52% lower), and lose on MiniMax M3 and on DeepSeek V4 Pro input. Two things change the sum in practice. Fireworks' Standard tier is shed first under load, so production traffic tends to end up on Priority or Fast at 20 to 50% more; we have one rate. And any model Fireworks has moved off serverless costs $8 per GPU-hour there, against cents per million tokens here.
For inference serving the gap with Nvidia is much narrower than it is for training, and on memory-bound workloads the MI325X's 256GB is an advantage: it sustains roughly 2.5x the GPT-OSS 120B throughput of an H100 on our infrastructure. Fireworks' Fast tier will beat our per-request speed on the handful of models it covers; if you need 100+ tokens a second on Kimi K3 today, that is a real advantage, priced at 1.5x. Where AMD matters more is model compatibility: most of the catalogue runs on it, not every model does. Ask before you migrate a specific model and we'll tell you straight. We also run Nvidia H100, B200 and B300 if that is the better fit.
Reserved, isolated GPU capacity plus the team to operate it, inference stack tuning, quantisation, batching, fine-tuning on your traffic, and on-call support. Not a separate professional services engagement. We benchmark your workload before you commit, then guarantee that performance level on the reserved capacity.
Yes, on qualifying models, applied automatically when a request reuses a prompt prefix of roughly 1,000+ tokens we served recently: Kimi K3 $0.30, GLM 5.2 $0.15, DeepSeek V4 Pro $0.20, DeepSeek V4 Flash $0.03, MiniMax M3 $0.06 and gpt-oss-120B $0.02 per million cached input tokens. Fireworks discounts more deeply on the DeepSeek models ($0.044 and $0.007), so if your workload is cache-heavy on DeepSeek, factor that in. Fireworks' cache also lives on a single replica and needs a session-affinity header to hit; ours needs no code change.
On Fireworks, the changelog announces the removal and points you at a successor; the last two rounds gave about two weeks' notice, and a model recommended as the migration target in June was itself removed in August. Our catalogue isn't curated for utilisation: public Hugging Face models with 100+ downloads are added automatically for supported architectures, less popular ones on request, and a model stays on the same API whether or not it is trending. If a model you depend on is on Fireworks' deprecation list, search our catalogue before you migrate. It is very likely already there.
Fine-tuning is available on the Business tier, on your dedicated capacity, and public full-weight fine-tunes of a supported architecture on Hugging Face are served at the base model's per-token rate (merge LoRAs first). We don't offer a discounted batch API, managed SFT or DPO with published training prices, GPU clusters by the hour, or Kimi K3 at its full 1M context; we serve 256K. If those are your main requirements, Fireworks is the better fit and we'd rather say so before you migrate.
Run the numbers on your own workload
Check whether your model is still on Fireworks' serverless list, then price it here in a minute, or take an API key and test on real traffic.