The Fireworks AI alternative that doesn't retire your model

Fireworks removes serverless models in waves: Llama 3.3 70B in May, gpt-oss-20B in August, GLM 5.2 and Kimi K2.6 on 25 September 2026. All four are still served here, next to 40,000+ other open models, on one OpenAI-compatible API at one rate.

A serverless tier that keeps getting smaller

Fireworks is fast and well engineered, and the right choice if you want managed fine-tuning or Blackwell GPUs by the hour.

◷

Models leave on two weeks' notice

Multiple rounds of serverless removals. The 25 September batch was announced on 12 September and takes GLM 5.2, Kimi K2.6, Kimi K2.7 Code and both DeepSeek V4 builds.

⊞

Three prices for one model

Standard is shed first under load. Priority costs 20 to 58% more, Fast costs 50% more. On Kimi K3 that is $15.00, $18.75 or $22.50 per million output tokens for the same weights.

↕

Your fine-tune needs its own GPU

LoRAs and uploaded models can't run on Fireworks serverless. They go on a dedicated deployment from $8 per GPU-hour, and a deployment scaled to zero answers 503 while it wakes.

◈

GPU prices rose in September

On 1 September 2026 Fireworks raised on-demand rates: H100 from $7 to $8 an hour, B200 from $10 to $13. Region-restricted deployments add another 50%.

Fireworks AI pricing vs ours, model by model

Fireworks' Standard serverless rates against ours, checked 16 September 2026. Cached input shown where Fireworks publishes it. Priority adds 20 to 58% and Fast adds 50% to every figure in the Fireworks column.

ModelFireworks
in / out
Featherless
in / out
Difference
Kimi K32780B MoE · 256K ctx$3.00 / $15.00$0.30 cached in$3.00 / $15.00level
GLM-5.3753B · 256K ctx$1.40 / $4.40$0.26 cached in$1.40 / $4.40level
GLM-5.2753B · off serverless 25 Sep 2026$1.40 / $4.40$0.14 cached in$1.40 / $4.40level
Kimi K2.61T MoE · off serverless 25 Sep 2026$0.95 / $4.00$0.16 cached in$0.80 / $3.40−15%
DeepSeek V4 Pro0813 build · off serverless 25 Sep 2026$1.32 / $3.96$0.044 cached in$1.60 / $3.20−9%
DeepSeek V4 Flash284B · 256K ctx$0.22 / $0.66$0.007 cached in$0.14 / $0.28−52%
MiniMax M3427B · 256K ctx$0.30 / $1.20$0.06 cached in$0.55 / $2.20+83%
gpt-oss-120B128K ctx$0.15 / $0.60$0.015 cached in$0.15 / $0.60level
Qwen3.5 397B A17BNot on serverlesson-demand GPU only$0.55 / $3.50—
Gemma 4 31BNot on serverlesson-demand GPU only$0.12 / $0.36—
gpt-oss-20B128K ctxRemoved from serverless27 Aug 2026$0.04 / $0.15—
Llama 3.3 70BRemoved from serverless14 May 2026$0.65 / $0.75—

What a deprecation costs

Fireworks removes Kimi K2.6 from serverless on 25 September 2026 and suggests GLM 5.3 or Kimi K3 instead. Say you want the model you shipped with. 100M input and 100M output tokens a month, list rates on both sides.

Fireworks — on-demand deployment
$43,800/mo

4 × B300 at $15.00 per GPU-hour for 730 hours. A 1T-parameter model at 8-bit needs about a terabyte of GPU memory, so the smallest deployment that holds Kimi K2.6 is four B300s, or eight H200s at $8.00 for $46,720. Scale-to-zero trims the bill but answers 503 while the replica warms, and a deployment with no traffic for 7 days is deleted.

Token cost on Featherless for Kimi K2.6
$420/mo

100M input at $0.80 = $80. 100M output at $3.40 = $340. The model never left the catalogue, so nothing in your integration changes.

$43,380 a month to keep a model that cost $495 on Fireworks serverless the week before. The gap is the product rather than a discount: serverless exists so a model doesn't need a node of its own. Fireworks' suggested replacements, GLM 5.3 and Kimi K3, are here too, at the same list rates Fireworks charges for them. Check your own model and volume rather than trusting either headline. Use the calculator →

Predictable pricing that scales with you

developer
Build production Al with the fastest usage-based inference.
$50 credits/month
  • Context size up to 256K
  • 1 agent environment included
  • Fastest response times
  • Unused credits roll over
  • Billed per token
Credit Amount
Subscribe
Billed monthly. Cancel anytime.
business
Dedicated GPUs and the team to run them.
Custom
  • Dedicated H100, MI325, B200 & B300 GPUs
  • Engineering team included
  • Gets cheaper over time with fine-tuning
  • Burst & failover to Public Cloud
Talk to an Engineer
Annual contracts. Volume pricing.

A managed node for a third of a Fireworks half node

Above a certain sustained volume, reserved capacity beats metered pricing on any platform. Fireworks publishes on-demand deployments at a per-GPU hourly rate, raised on 1 September 2026, so the comparison is arithmetic.

Fireworks — half node, on demand
$23,360/mo

4 × H100 at $8.00 per GPU per hour, running continuously (730 hrs). A full 8-GPU node is $46,720. On B200 at $13.00/hr the same half node is $37,960, and a region-restricted deployment adds 50% on top.

Featherless — full managed node
$7,500/mo

A dedicated AMD MI325X node, reserved and isolated. Inference stack tuning, quantisation, batching and on-call support are included rather than billed separately.

How it works in practice: Available in the US, EU and Southeast Asia with no procurement cycle.

Switching from Fireworks AI: a one-line change

Both APIs are OpenAI-compatible. If you're calling Fireworks through the OpenAI SDK today, swap the base URL and key, then replace the accounts/fireworks/models/ prefix with the Hugging Face repo path.

Get your API key
# Before — Fireworks AI client = OpenAI( base_url="https://api.fireworks.ai/inference/v1", api_key=os.environ["FIREWORKS_API_KEY"] )   # After — Featherless client = OpenAI( base_url="https://api.featherless.ai/v1", api_key=os.environ["FEATHERLESS_API_KEY"] )   # accounts/fireworks/models/glm-5p2 becomes the Hugging Face repo path r = client.chat.completions.create( model="zai-org/GLM-5.2", messages=[{"role": "user", "content": "Hello"}] )

Frequently asked questions

What is the best Fireworks AI alternative?

For serverless open-model inference, Featherless is the closest like-for-like: the same OpenAI-compatible API, the same flagship models at level or lower list rates (Kimi K3 $3.00 / $15.00 on both; DeepSeek V4 Flash $0.42 against $0.88 per 1M in + 1M out, September 2026), plus 40,000+ open models that Fireworks serves only on rented GPUs or not at all. Fireworks remains the better choice for managed fine-tuning, batch inference at half price, speculative decoding on dedicated deployments, and Blackwell GPUs by the hour.

Can I use the $25 Chat plan for my application?

No. Chat is sold for interactive, human-driven use by the subscriber, not for API traffic, background automation, reselling or benchmarking. If an application is making the calls, you want the Developer plan. We're explicit about this because signing up for the wrong plan is the single most common source of frustration when people switch.

Is Featherless cheaper than Fireworks?

On list rates, mostly level. We match Fireworks Standard on Kimi K3, GLM 5.2, GLM 5.3 and gpt-oss-120B, undercut it on Kimi K2.6 (15% lower) and DeepSeek V4 Flash (52% lower), and lose on MiniMax M3 and on DeepSeek V4 Pro input. Two things change the sum in practice. Fireworks' Standard tier is shed first under load, so production traffic tends to end up on Priority or Fast at 20 to 50% more; we have one rate. And any model Fireworks has moved off serverless costs $8 per GPU-hour there, against cents per million tokens here.

Is AMD hardware slower for inference?

For inference serving the gap with Nvidia is much narrower than it is for training, and on memory-bound workloads the MI325X's 256GB is an advantage: it sustains roughly 2.5x the GPT-OSS 120B throughput of an H100 on our infrastructure. Fireworks' Fast tier will beat our per-request speed on the handful of models it covers; if you need 100+ tokens a second on Kimi K3 today, that is a real advantage, priced at 1.5x. Where AMD matters more is model compatibility: most of the catalogue runs on it, not every model does. Ask before you migrate a specific model and we'll tell you straight. We also run Nvidia H100, B200 and B300 if that is the better fit.

What does the $7,500/month dedicated node include?

Reserved, isolated GPU capacity plus the team to operate it, inference stack tuning, quantisation, batching, fine-tuning on your traffic, and on-call support. Not a separate professional services engagement. We benchmark your workload before you commit, then guarantee that performance level on the reserved capacity.

Do you offer cached input pricing?

Yes, on qualifying models, applied automatically when a request reuses a prompt prefix of roughly 1,000+ tokens we served recently: Kimi K3 $0.30, GLM 5.2 $0.15, DeepSeek V4 Pro $0.20, DeepSeek V4 Flash $0.03, MiniMax M3 $0.06 and gpt-oss-120B $0.02 per million cached input tokens. Fireworks discounts more deeply on the DeepSeek models ($0.044 and $0.007), so if your workload is cache-heavy on DeepSeek, factor that in. Fireworks' cache also lives on a single replica and needs a session-affinity header to hit; ours needs no code change.

What happens when a model I use is deprecated?

On Fireworks, the changelog announces the removal and points you at a successor; the last two rounds gave about two weeks' notice, and a model recommended as the migration target in June was itself removed in August. Our catalogue isn't curated for utilisation: public Hugging Face models with 100+ downloads are added automatically for supported architectures, less popular ones on request, and a model stays on the same API whether or not it is trending. If a model you depend on is on Fireworks' deprecation list, search our catalogue before you migrate. It is very likely already there.

Do you offer fine-tuning, batch inference or reserved throughput?

Fine-tuning is available on the Business tier, on your dedicated capacity, and public full-weight fine-tunes of a supported architecture on Hugging Face are served at the base model's per-token rate (merge LoRAs first). We don't offer a discounted batch API, managed SFT or DPO with published training prices, GPU clusters by the hour, or Kimi K3 at its full 1M context; we serve 256K. If those are your main requirements, Fireworks is the better fit and we'd rather say so before you migrate.

Run the numbers on your own workload

Check whether your model is still on Fireworks' serverless list, then price it here in a minute, or take an API key and test on real traffic.