Cerebras
Cerebras serves inference from wafer-scale WSE-3 parts and is repeatedly benchmarked as the fastest hosted inference available, at 1,800 to 3,000-plus tokens per second on open models. The company frames this as up to 30 times faster than GPU systems. The question a buyer needs answered is what that speed costs relative to the identical model elsewhere, and that number is knowable.
| Category | Custom-silicon speed specialists |
|---|---|
| Region | US |
| What it serves | Llama, Qwen, DeepSeek distills, GPT-OSS. No proprietary frontier models and no custom-model uploads. |
| Pricing model | Per token |
| Representative pricing | ~$0.10-6 per Mtok depending on model ($0.35/M cheapest input). Pay-as-you-go and enterprise tiers. |
| OpenAI-compatible API | Yes |
| Free tier | Free trial: $5 in credits, all models. |
| Official pages | Site · Pricing · Docs |
What Cerebras serves
An open-weight catalog: Llama, Qwen including Qwen3 235B and Qwen3 Coder 480B, DeepSeek distills and GPT-OSS. There are no proprietary frontier models, and custom model weights and fine-tuning are enterprise-tier features rather than self-serve ones. Serving is from US infrastructure.
How Cerebras pricing works
Three tiers. A free trial gives $5 in credits after account creation with access to all models. A self-serve Developer tier starts at $10 and carries ten times the free rate limits plus elevated processing priority. Enterprise adds the highest limits, custom weight support, fine-tuning and dedicated support. Per-token rates run roughly $0.10 to $6 per million tokens depending on model. The comparison that matters: on GPT-OSS-120B, Cerebras lists $0.35 input and $0.75 output per million tokens, against $0.15 and $0.60 at Groq, Together, Nebius, Amazon Bedrock and DeepInfra for the identical weights, and $0.039 and $0.10 at FlexAI. So Cerebras is about 2.3 times the common input rate and roughly 9 times the cheapest listing, for the same model. That is the wafer-scale premium stated as a number rather than as a claim.
To compare this against the rest of the market rather than reading it in isolation, see the value leaderboard and the coding cost calculator.
Limitations and things to check first
The free tier is $5 in credits, not the 1,000,000 tokens per day figure that circulates widely in secondary coverage of Cerebras. Check the pricing page rather than a roundup before budgeting against it. The catalog is open-weight only, so there is no path to a frontier proprietary model. Custom weights and fine-tuning sit behind the enterprise tier, which makes Cerebras a poor fit for a team whose differentiation is its own tuned model unless it is willing to have that conversation. And the premium is real: on any workload where throughput rather than first-token latency is the binding constraint, the same model is available for well under half the price elsewhere.
Who Cerebras suits
Workloads where latency is the product and the model is already an open-weight one: real-time voice, agentic loops that chain many short calls, and interactive tools where a two-second wait loses the user. It is a poor fit for batch or offline work, where nobody is waiting and the 2.3-times input premium over the same weights elsewhere buys nothing, and for teams that need custom fine-tuned weights on a self-serve plan.
Providers to weigh against this one
Groq
Custom-silicon (LPU) inference cloud for open models; among the fastest inference available.
SambaNova
Custom-silicon (RDU) inference platform (SambaCloud) purpose-built for agentic inference.
Cerebras: frequently asked questions
- How much does Cerebras inference cost?
- Roughly $0.10 to $6 per million tokens depending on the model, billed per token. As a concrete anchor, GPT-OSS-120B is listed at $0.35 input and $0.75 output per million tokens. A self-serve Developer tier starts at $10 and an enterprise tier is quoted separately.
- Is Cerebras more expensive than Groq?
- On the same model, yes. GPT-OSS-120B is $0.35 input and $0.75 output per million tokens on Cerebras against $0.15 and $0.60 on Groq, so about 2.3 times the input rate for identical weights. Cerebras is generally the faster of the two, so the premium buys tokens per second, not capability.
- Does Cerebras have a free tier?
- It offers a free trial of $5 in credits after you create an account, with access to all models and community Discord support. Secondary coverage often reports a 1,000,000 token per day free allowance, which does not match the current pricing page.
- Can you fine-tune a model on Cerebras?
- Not on the self-serve tiers. Custom model weight support and fine-tuning services are listed as enterprise-tier features, alongside the highest rate limits and a dedicated support team. The free and Developer tiers serve the stock open-weight catalog only.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.