Groq
Groq sells inference on its own LPU silicon rather than on GPUs, and the whole proposition is latency: 500 to 1,000-plus tokens per second on a catalog that is deliberately open-weight only. It is a speed vendor first and a price vendor second, and on the models where those two things can be compared directly, the pricing is more interesting than the marketing suggests.
| Category | Custom-silicon speed specialists |
|---|---|
| Region | US |
| What it serves | Runs Llama, Qwen, Kimi, GPT-OSS, DeepSeek distills and Whisper at 500-1,000+ tok/s. Catalog is open-source only. |
| Pricing model | Per token |
| Representative pricing | Llama 3.1 8B $0.05/$0.08, Llama 3.3 70B $0.59/$0.79, Kimi K2 $1/$3 per Mtok. Batch and caching each cut 50%. |
| OpenAI-compatible API | Yes |
| Free tier | Free developer tier (no credit card). |
| Official pages | Site · Pricing · Docs |
What Groq serves
The catalog is open-source models only: Llama, Qwen, Kimi, GPT-OSS and DeepSeek distills for text, plus Whisper for speech. There is no proprietary frontier model here, and no facility to upload your own weights. Groq is a place to run models someone else trained, very fast, from a US region.
How Groq pricing works
Billing is per token with no subscription, and the rate depends entirely on which model you call: roughly $0.05 input and $0.08 output per million tokens for Llama 3.1 8B, up to $1 and $3 for Kimi K2. Batch submission and prompt caching each cut the rate by half and stack, so a cached batch workload pays about a quarter of the on-demand price. The number worth knowing is comparative. On GPT-OSS-120B, Groq lists $0.15 input and $0.60 output per million tokens, which is exactly what Together, Nebius, Amazon Bedrock and DeepInfra charge for the identical weights, with cached input reads at $0.075. The speed is not billed as a premium on that row: it is the common market rate with a faster machine behind it.
To compare this against the rest of the market rather than reading it in isolation, see the value leaderboard and the coding cost calculator.
Limitations and things to check first
Free-plan limits are set per model rather than per account, and they are tight: several models cap at 30 requests a minute, 1,000 a day, 8,000 tokens a minute and 200,000 tokens a day, and the documentation says the authoritative figures are the ones in your own account settings rather than the published table. Higher limits are described as available for select workloads and enterprise use. The open-weight-only catalog is the hard ceiling: no frontier proprietary model, no custom weights. One structural item belongs on any Groq decision: NVIDIA agreed to pay about $20 billion for a perpetual license to Groq LPU patents, finalized on December 24 2025. GroqCloud continues to operate, but a platform whose core silicon IP has been licensed to the dominant GPU vendor carries a roadmap question that a pricing page does not answer.
Who Groq suits
Latency-bound work on open models: voice agents, interactive coding assistants, search and rerank paths, anything where a person is waiting on the first token. It suits a team that has already chosen an open-weight model and wants it served faster than a GPU host will serve it. It does not suit a team that needs a frontier proprietary model, custom fine-tuned weights, or one vendor for both.
Providers to weigh against this one
Cerebras
Wafer-scale (WSE-3) inference cloud; the speed champion at 1,800-3,000+ tok/s on open models.
SambaNova
Custom-silicon (RDU) inference platform (SambaCloud) purpose-built for agentic inference.
Groq: frequently asked questions
- How much does Groq cost per million tokens?
- It depends on the model, and there is no subscription. Rates run from about $0.05 input and $0.08 output per million tokens on Llama 3.1 8B to $1 and $3 on Kimi K2, with GPT-OSS-120B at $0.15 and $0.60. Batch submission and prompt caching each halve the rate and stack, so a cached batch job pays roughly a quarter of on-demand.
- Is Groq cheaper than other inference providers?
- Not systematically, no. On GPT-OSS-120B its $0.15 input rate matches Together, Nebius, Amazon Bedrock and DeepInfra to the cent for identical weights, while FlexAI lists the same model at $0.039. Groq competes on tokens per second, not on being the cheapest place to run a given model.
- Does Groq have a free tier?
- Yes, a free developer plan with no credit card required. The limits are applied per model and are small, on the order of 30 requests a minute and 200,000 tokens a day on several models, so it is sized for evaluation rather than for serving traffic.
- Can you run GPT-5 or Claude on Groq?
- No. The catalog is open-weight models only, so Llama, Qwen, Kimi, GPT-OSS and DeepSeek distills are available and the proprietary frontier models from OpenAI, Anthropic and Google are not. GPT-OSS is OpenAI open-weight release, not the GPT-5 family.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.