The Replicate alternative for open-model LLM inference

Featherless serves Kimi K3, DeepSeek V4, GLM 5.3, Gemma 4 and 40,000+ other open models on one OpenAI-compatible API, billed per token.

Built for images, borrowed for LLMs

Replicate is the best place on the internet to run FLUX, Seedream, Veo or Kling behind an API, and since December 2025 it has been part of Cloudflare. Language models were never the point. Its LLM catalogue mixes proprietary APIs with open models from 2024 and 2025, and the new open-weight releases Cloudflare ships go to Workers AI, not to Replicate.

◷

The shelf is stale

No Kimi K3, no DeepSeek V4, no GLM 5, no MiniMax M3. The newest official DeepSeek is V3.1; the newest open-weight Qwen is the July 2025 Qwen3 235B. Gemma 4 31B exists only as a community push, quantised to 4-bit on a single L40S.

⊞

No chat completions endpoint

Replicate's API is predictions: create, then poll or stream. Its open LLMs take a flat prompt and system_prompt, return an array of string chunks, and expose no messages array or tools parameter. An OpenAI-SDK integration needs a shim.

↕

Cold boots outside the official list

Official models stay warm. Community models, which is where Gemma 4 and most recent open LLMs live, boot when called, and Replicate's own docs say that in some cases this can take several minutes.

◈

Per-second billing for the rest

Anything outside the official list bills by the hardware-second: H100 $5.49/hr, A100 80GB $5.04/hr, L40S $3.51/hr. Deployments bill for setup, idle and active time, failed runs included.

Replicate pricing vs ours, model by model

Replicate's per-token rates as displayed on its model pages on 16 September 2026, against ours. The line under each Replicate price says whether the model is official (kept warm) or a community push (billed by the GPU-second, with cold boots).

ModelReplicate
in / out
Featherless
in / out
Difference
Kimi K2.6256K ctx · Cloudflare pass-through$0.95 / $4.00official · Workers AI rate$0.80 / $3.40−15%
DeepSeek V3.1685B · 128K ctx$0.672 / $2.016official model$0.285 / $1.00−52%
Qwen3 235B A22BInstruct-2507 build on Replicate$0.264 / $1.06official model$0.455 / $1.82+72%
gpt-oss-120B128K ctx$0.18 / $0.72official model$0.15 / $0.60−17%
gpt-oss-20B128K ctx$0.09 / $0.36official model$0.04 / $0.15−58%
Gemma 4 31Bcommunity push · 4-bit on one L40S$3.51 / hrbilled per GPU-second$0.12 / $0.36—
GLM-5.3
753B · 256K ctx
Not on Replicateno GLM 5 at all$1.40 / $4.40—
Kimi K3
2.78T · 256K ctx
Not on Replicatenewest Kimi is K2.6$3.00 / $15.00—
DeepSeek V4 Pro
1.6T · 256K ctx
Not on Replicatenewest DeepSeek is V3.1$1.60 / $3.20—
DeepSeek V4.1 Flash763B · 256K ctxNot on Replicate$0.30 / $1.20—
Qwen3.8 Flash Next180B · 256K ctxNot on Replicatenewest open Qwen is Jul 2025$0.15 / $0.50—
MiniMax M3256K ctx · Replicate hosts MiniMax video, not M-seriesNot on Replicate$0.55 / $2.20—

Pay only for tokens you consume

Replicate — deployment, one 8-GPU node
$32,062/mo

8 × H200 at $43.92 per hour, running continuously (730 hrs), the smallest listed node that holds a 753B model at 8-bit. Replicate sells that tier only with a committed-spend contract, and bills deployments for setup, idle and active time. Scale to zero and the first request after a quiet spell waits for the boot.

Token cost on Featherless for GLM-5.3
$580/mo

100M input at $1.40 = $140. 100M output at $4.40 = $440. Nothing to push and nothing to keep warm; the model is already in the catalogue under its Hugging Face repo path.

Predictable pricing that scales with you

developer
Build production Al with the fastest usage-based inference.
$50 credits/month
  • Context size up to 256K
  • 1 agent environment included
  • Fastest response times
  • Unused credits roll over
  • Billed per token
Credit Amount
Subscribe
Billed monthly. Cancel anytime.
business
Dedicated GPUs and the team to run them.
Custom
  • Dedicated H100, MI325, B200 & B300 GPUs
  • Engineering team included
  • Gets cheaper over time with fine-tuning
  • Burst & failover to Public Cloud
Talk to an Engineer
Annual contracts. Volume pricing.

A managed node for less than half a node elsewhere

Above a certain sustained volume, reserved capacity beats metered pricing on any platform. Replicate publishes multi-GPU hardware at a per-second rate, so the comparison is arithmetic, with one catch: 2×, 4× and 8× H100 instances are only sold under committed-spend contracts.

Replicate — half node, committed spend
$16,031/mo

4 × H100 at $5.49 per GPU per hour ($0.0061 a second), running continuously (730 hrs). A full 8-GPU node is $32,062. Replicate lists no B200 and no AMD hardware; H200 is priced the same as H100.

Featherless — full managed node
$7,500/mo

A dedicated AMD MI325X node, reserved and isolated. Inference stack tuning, quantisation, batching and on-call support are included rather than billed separately.

How it works in practice: Available in the US, EU and Southeast Asia with no procurement cycle.

Switching from Replicate: swap the SDK, keep the prompt

Replicate has no OpenAI-compatible endpoint, so this is a small rewrite rather than a base-URL change. The OpenAI SDK replaces the replicate client, and system_prompt plus prompt become a messages array. Streaming comes back on the same call with stream=True.

Get your API key
# Before — Replicate import replicate out = replicate.run( "moonshotai/kimi-k2.6", input={"system_prompt": "You are terse.", "prompt": "Hello"} ) print("".join(out))   # After — Featherless client = OpenAI( base_url="https://api.featherless.ai/v1", api_key=os.environ["FEATHERLESS_API_KEY"] ) r = client.chat.completions.create( model="moonshotai/Kimi-K2.6", messages=[{"role": "system", "content": "You are terse."}, {"role": "user", "content": "Hello"}] )

Frequently asked questions

What is the best Replicate alternative?

For language models, Featherless: 40,000+ open models on an OpenAI-compatible API, billed per token, including the 2026 releases Replicate doesn't carry (Kimi K3, DeepSeek V4, GLM 5.3, Qwen3.8, MiniMax M3). Where Replicate does serve a current open LLM, our rates are lower on gpt-oss-120B (17%), gpt-oss-20B (58%), DeepSeek V3.1 (52%) and Kimi K2.6 (15%). For image, video and audio generation, or for pushing your own container, Replicate is the better tool and we don't compete with it.

Can I use the $25 Chat plan for my application?

No. Chat is sold for interactive, human-driven use by the subscriber, not for API traffic, background automation, reselling or benchmarking. If an application is making the calls, you want the Developer plan. We're explicit about this because signing up for the wrong plan is the single most common source of frustration when people switch.

What changed at Replicate after the Cloudflare acquisition?

Cloudflare announced the acquisition on 17 November 2025 and closed it on 1 December; by April 2026 the Replicate team had been folded into Cloudflare's AI Platform group. The API and pricing were kept, and the status page moved to cloudflarestatus.com in September 2026. For LLM users the practical change is where new models land: DeepSeek V4, GLM 5.2 and 5.3, Kimi K2.6 and K2.7, Gemma 4 and Qwen 3.8 all shipped on Workers AI in 2026, and Replicate's own changelog has been quiet since April. Kimi K2.6 on Replicate is a pass-through to Workers AI at Cloudflare's rate.

Is AMD hardware slower for inference?

For inference serving the gap with Nvidia is much narrower than it is for training, and on memory-bound workloads the MI325X's 256GB is an advantage: it sustains roughly 2.5x the GPT-OSS 120B throughput of an H100 on our infrastructure. Replicate's hardware list is Nvidia only, from T4 to H200. Where AMD matters is model compatibility: most of the catalogue runs on it, not every model does. Ask before you migrate a specific model and we'll tell you straight. We also run Nvidia H100, B200 and B300 if that is the better fit.

What does the $7,500/month dedicated node include?

Reserved, isolated GPU capacity plus the team to operate it, inference stack tuning, quantisation, batching, fine-tuning on your traffic, and on-call support. Not a separate professional services engagement. We benchmark your workload before you commit, then guarantee that performance level on the reserved capacity.

Can I push my own model, like Cog on Replicate?

Not as a container. We serve public Hugging Face models of supported architectures: models with 100+ downloads are added automatically, others on request, and full-weight fine-tunes of a supported base run at the base model's rate. Business customers deploy models from their dashboard. If your model is a custom architecture, needs its own pre- or post-processing code, or isn't a language model, Cog on Replicate is the right tool. Replicate also serves Llama 4 Maverick and Scout by the token; search our catalogue for the exact models you rely on before you move.

Do you offer cached input pricing?

Yes, on qualifying models, applied automatically when a request reuses a prompt prefix of roughly 1,000+ tokens we served recently: Kimi K3 $0.30, Kimi K2.6 $0.154, GLM 5.2 $0.15 and gpt-oss-120B $0.02 per million cached input tokens. Replicate publishes no cached-input rate on its LLM pages.

Do you serve image, video or audio models?

No. Featherless is text and vision-language models only. If your stack is FLUX, Seedream, Wan or ElevenLabs on Replicate, keep it there and point only your text traffic at us. We don't rent GPUs by the second either; the closest we come is the managed dedicated node above.

Run the numbers on your own workload

Price your text traffic against both in a minute, or take an API key and test on real traffic.