Capital & Compute

Cerebras

Inference provider· Custom-silicon speed specialists· Checked 2026-09-01

Cerebras serves inference from wafer-scale WSE-3 parts and is repeatedly benchmarked as the fastest hosted inference available, at 1,800 to 3,000-plus tokens per second on open models. The company frames this as up to 30 times faster than GPU systems. The question a buyer needs answered is what that speed costs relative to the identical model elsewhere, and that number is knowable.

Key facts about Cerebras
CategoryCustom-silicon speed specialists
RegionUS
What it servesLlama, Qwen, DeepSeek distills, GPT-OSS. No proprietary frontier models and no custom-model uploads.
Pricing modelPer token
Representative pricing~$0.10-6 per Mtok depending on model ($0.35/M cheapest input). Pay-as-you-go and enterprise tiers.
OpenAI-compatible APIYes
Free tierFree trial: $5 in credits, all models.
Official pagesSite · Pricing · Docs

What Cerebras serves

An open-weight catalog: Llama, Qwen including Qwen3 235B and Qwen3 Coder 480B, DeepSeek distills and GPT-OSS. There are no proprietary frontier models, and custom model weights and fine-tuning are enterprise-tier features rather than self-serve ones. Serving is from US infrastructure.

How Cerebras pricing works

Three tiers. A free trial gives $5 in credits after account creation with access to all models. A self-serve Developer tier starts at $10 and carries ten times the free rate limits plus elevated processing priority. Enterprise adds the highest limits, custom weight support, fine-tuning and dedicated support. Per-token rates run roughly $0.10 to $6 per million tokens depending on model. The comparison that matters: on GPT-OSS-120B, Cerebras lists $0.35 input and $0.75 output per million tokens, against $0.15 and $0.60 at Groq, Together, Nebius, Amazon Bedrock and DeepInfra for the identical weights, and $0.039 and $0.10 at FlexAI. So Cerebras is about 2.3 times the common input rate and roughly 9 times the cheapest listing, for the same model. That is the wafer-scale premium stated as a number rather than as a claim.

To compare this against the rest of the market rather than reading it in isolation, see the value leaderboard and the coding cost calculator.

Limitations and things to check first

The free tier is $5 in credits, not the 1,000,000 tokens per day figure that circulates widely in secondary coverage of Cerebras. Check the pricing page rather than a roundup before budgeting against it. The catalog is open-weight only, so there is no path to a frontier proprietary model. Custom weights and fine-tuning sit behind the enterprise tier, which makes Cerebras a poor fit for a team whose differentiation is its own tuned model unless it is willing to have that conversation. And the premium is real: on any workload where throughput rather than first-token latency is the binding constraint, the same model is available for well under half the price elsewhere.

Who Cerebras suits

Workloads where latency is the product and the model is already an open-weight one: real-time voice, agentic loops that chain many short calls, and interactive tools where a two-second wait loses the user. It is a poor fit for batch or offline work, where nobody is waiting and the 2.3-times input premium over the same weights elsewhere buys nothing, and for teams that need custom fine-tuned weights on a self-serve plan.

Providers to weigh against this one

Cerebras: frequently asked questions

How much does Cerebras inference cost?
Roughly $0.10 to $6 per million tokens depending on the model, billed per token. As a concrete anchor, GPT-OSS-120B is listed at $0.35 input and $0.75 output per million tokens. A self-serve Developer tier starts at $10 and an enterprise tier is quoted separately.
Is Cerebras more expensive than Groq?
On the same model, yes. GPT-OSS-120B is $0.35 input and $0.75 output per million tokens on Cerebras against $0.15 and $0.60 on Groq, so about 2.3 times the input rate for identical weights. Cerebras is generally the faster of the two, so the premium buys tokens per second, not capability.
Does Cerebras have a free tier?
It offers a free trial of $5 in credits after you create an account, with access to all models and community Discord support. Secondary coverage often reports a 1,000,000 token per day free allowance, which does not match the current pricing page.
Can you fine-tune a model on Cerebras?
Not on the self-serve tiers. Custom model weight support and fine-tuning services are listed as enterprise-tier features, alongside the highest rate limits and a dedicated support team. The free and Developer tiers serve the stock open-weight catalog only.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 35 providers in the directory