The Replicate alternative for open-model LLM inference
Featherless serves Kimi K3, DeepSeek V4, GLM 5.3, Gemma 4 and 40,000+ other open models on one OpenAI-compatible API, billed per token.
Built for images, borrowed for LLMs
Replicate is the best place on the internet to run FLUX, Seedream, Veo or Kling behind an API, and since December 2025 it has been part of Cloudflare. Language models were never the point. Its LLM catalogue mixes proprietary APIs with open models from 2024 and 2025, and the new open-weight releases Cloudflare ships go to Workers AI, not to Replicate.
The shelf is stale
No Kimi K3, no DeepSeek V4, no GLM 5, no MiniMax M3. The newest official DeepSeek is V3.1; the newest open-weight Qwen is the July 2025 Qwen3 235B. Gemma 4 31B exists only as a community push, quantised to 4-bit on a single L40S.
No chat completions endpoint
Replicate's API is predictions: create, then poll or stream. Its open LLMs take a flat prompt and system_prompt, return an array of string chunks, and expose no messages array or tools parameter. An OpenAI-SDK integration needs a shim.
Cold boots outside the official list
Official models stay warm. Community models, which is where Gemma 4 and most recent open LLMs live, boot when called, and Replicate's own docs say that in some cases this can take several minutes.
Per-second billing for the rest
Anything outside the official list bills by the hardware-second: H100 $5.49/hr, A100 80GB $5.04/hr, L40S $3.51/hr. Deployments bill for setup, idle and active time, failed runs included.
Replicate pricing vs ours, model by model
Replicate's per-token rates as displayed on its model pages on 16 September 2026, against ours. The line under each Replicate price says whether the model is official (kept warm) or a community push (billed by the GPU-second, with cold boots).
| Model | Replicate in / out | Featherless in / out | Difference |
|---|---|---|---|
| Kimi K2.6256K ctx · Cloudflare pass-through | $0.95 / $4.00official · Workers AI rate | $0.80 / $3.40 | −15% |
| DeepSeek V3.1685B · 128K ctx | $0.672 / $2.016official model | $0.285 / $1.00 | −52% |
| Qwen3 235B A22BInstruct-2507 build on Replicate | $0.264 / $1.06official model | $0.455 / $1.82 | +72% |
| gpt-oss-120B128K ctx | $0.18 / $0.72official model | $0.15 / $0.60 | −17% |
| gpt-oss-20B128K ctx | $0.09 / $0.36official model | $0.04 / $0.15 | −58% |
| Gemma 4 31Bcommunity push · 4-bit on one L40S | $3.51 / hrbilled per GPU-second | $0.12 / $0.36 | — |
| GLM-5.3 753B · 256K ctx | Not on Replicateno GLM 5 at all | $1.40 / $4.40 | — |
| Kimi K3 2.78T · 256K ctx | Not on Replicatenewest Kimi is K2.6 | $3.00 / $15.00 | — |
| DeepSeek V4 Pro 1.6T · 256K ctx | Not on Replicatenewest DeepSeek is V3.1 | $1.60 / $3.20 | — |
| DeepSeek V4.1 Flash763B · 256K ctx | Not on Replicate | $0.30 / $1.20 | — |
| Qwen3.8 Flash Next180B · 256K ctx | Not on Replicatenewest open Qwen is Jul 2025 | $0.15 / $0.50 | — |
| MiniMax M3256K ctx · Replicate hosts MiniMax video, not M-series | Not on Replicate | $0.55 / $2.20 | — |
Pay only for tokens you consume
8 × H200 at $43.92 per hour, running continuously (730 hrs), the smallest listed node that holds a 753B model at 8-bit. Replicate sells that tier only with a committed-spend contract, and bills deployments for setup, idle and active time. Scale to zero and the first request after a quiet spell waits for the boot.
100M input at $1.40 = $140. 100M output at $4.40 = $440. Nothing to push and nothing to keep warm; the model is already in the catalogue under its Hugging Face repo path.
Predictable pricing that scales with you
- Context size up to 256K
- 1 agent environment included
- Fastest response times
- Unused credits roll over
- Billed per token
- Dedicated H100, MI325, B200 & B300 GPUs
- Engineering team included
- Gets cheaper over time with fine-tuning
- Burst & failover to Public Cloud
A managed node for less than half a node elsewhere
Above a certain sustained volume, reserved capacity beats metered pricing on any platform. Replicate publishes multi-GPU hardware at a per-second rate, so the comparison is arithmetic, with one catch: 2×, 4× and 8× H100 instances are only sold under committed-spend contracts.
4 × H100 at $5.49 per GPU per hour ($0.0061 a second), running continuously (730 hrs). A full 8-GPU node is $32,062. Replicate lists no B200 and no AMD hardware; H200 is priced the same as H100.
A dedicated AMD MI325X node, reserved and isolated. Inference stack tuning, quantisation, batching and on-call support are included rather than billed separately.
Switching from Replicate: swap the SDK, keep the prompt
Replicate has no OpenAI-compatible endpoint, so this is a small rewrite rather than a base-URL change. The OpenAI SDK replaces the replicate client, and system_prompt plus prompt become a messages array. Streaming comes back on the same call with stream=True.
Get your API keyFrequently asked questions
For language models, Featherless: 40,000+ open models on an OpenAI-compatible API, billed per token, including the 2026 releases Replicate doesn't carry (Kimi K3, DeepSeek V4, GLM 5.3, Qwen3.8, MiniMax M3). Where Replicate does serve a current open LLM, our rates are lower on gpt-oss-120B (17%), gpt-oss-20B (58%), DeepSeek V3.1 (52%) and Kimi K2.6 (15%). For image, video and audio generation, or for pushing your own container, Replicate is the better tool and we don't compete with it.
No. Chat is sold for interactive, human-driven use by the subscriber, not for API traffic, background automation, reselling or benchmarking. If an application is making the calls, you want the Developer plan. We're explicit about this because signing up for the wrong plan is the single most common source of frustration when people switch.
Cloudflare announced the acquisition on 17 November 2025 and closed it on 1 December; by April 2026 the Replicate team had been folded into Cloudflare's AI Platform group. The API and pricing were kept, and the status page moved to cloudflarestatus.com in September 2026. For LLM users the practical change is where new models land: DeepSeek V4, GLM 5.2 and 5.3, Kimi K2.6 and K2.7, Gemma 4 and Qwen 3.8 all shipped on Workers AI in 2026, and Replicate's own changelog has been quiet since April. Kimi K2.6 on Replicate is a pass-through to Workers AI at Cloudflare's rate.
For inference serving the gap with Nvidia is much narrower than it is for training, and on memory-bound workloads the MI325X's 256GB is an advantage: it sustains roughly 2.5x the GPT-OSS 120B throughput of an H100 on our infrastructure. Replicate's hardware list is Nvidia only, from T4 to H200. Where AMD matters is model compatibility: most of the catalogue runs on it, not every model does. Ask before you migrate a specific model and we'll tell you straight. We also run Nvidia H100, B200 and B300 if that is the better fit.
Reserved, isolated GPU capacity plus the team to operate it, inference stack tuning, quantisation, batching, fine-tuning on your traffic, and on-call support. Not a separate professional services engagement. We benchmark your workload before you commit, then guarantee that performance level on the reserved capacity.
Not as a container. We serve public Hugging Face models of supported architectures: models with 100+ downloads are added automatically, others on request, and full-weight fine-tunes of a supported base run at the base model's rate. Business customers deploy models from their dashboard. If your model is a custom architecture, needs its own pre- or post-processing code, or isn't a language model, Cog on Replicate is the right tool. Replicate also serves Llama 4 Maverick and Scout by the token; search our catalogue for the exact models you rely on before you move.
Yes, on qualifying models, applied automatically when a request reuses a prompt prefix of roughly 1,000+ tokens we served recently: Kimi K3 $0.30, Kimi K2.6 $0.154, GLM 5.2 $0.15 and gpt-oss-120B $0.02 per million cached input tokens. Replicate publishes no cached-input rate on its LLM pages.
No. Featherless is text and vision-language models only. If your stack is FLUX, Seedream, Wan or ElevenLabs on Replicate, keep it there and point only your text traffic at us. We don't rent GPUs by the second either; the closest we come is the managed dedicated node above.
Run the numbers on your own workload
Price your text traffic against both in a minute, or take an API key and test on real traffic.