Grok 4.6 Agent Cost: The 200K Price Cliff
Grok 4.6 kept the $2 and $6 headline but raised cached input 67 percent, and doubles every rate past 200K tokens. What that does to an agent bill.
SpaceXAI shipped Grok 4.6 on August 12, 2026, thirty-five days after Grok 4.5, and the coverage settled on one number: $2 per million input tokens and $6 per million output, unchanged from its predecessor, against $5 and $25 for Claude Opus 5. Same price, more intelligence. That is true of the headline and false of the bill, because the model is pitched at long-running agents and the two rate-card lines that dominate a long-running agent’s bill both went up.
What actually changed on the rate card
Grok 4.6 is documented by xAI with a 500,000-token context window, text and image input, and a February 1, 2026 knowledge cutoff. It is available through the xAI API, Cursor, Grok Build, OpenRouter, Vercel and Cloudflare, and the launch announcement describes it as built “for long-running agents and more ambitious interactive and visual work.”
Hold that phrase, because the pricing has to be read against it. Here is the full card, next to the model it replaced.
| Line item | Grok 4.5 | Grok 4.6 | Change |
|---|---|---|---|
| Input, under 200K | $2.00 | $2.00 | unchanged |
| Cached input, under 200K | $0.30 | $0.50 | +67% |
| Output, under 200K | $6.00 | $6.00 | unchanged |
| Input, 200K and above | $4.00 | $4.00 | unchanged |
| Cached input, 200K and above | $0.60 | $1.00 | +67% |
| Output, 200K and above | $12.00 | $12.00 | unchanged |
Every rate held except one, and the one that moved is the one a long-running agent spends most of its money on. An agent loop resends its accumulated context on every turn. After the first call, the overwhelming majority of the tokens it bills for are cache reads, not fresh input and not output. Raising the cache-read rate by two thirds while leaving the headline untouched is a price rise that does not look like one.
The 200K cliff is a step, not a slope
The second thing the rate card does is unusual enough to be worth stating precisely. xAI’s documentation says long-context pricing applies once prompts exceed 200,000 tokens, “billed at the higher rate for all tokens.”
All tokens. Not the tokens above the line. A request with a 199,000-token prompt bills at $2 / $0.50 / $6. A request with a 201,000-token prompt bills at $4 / $1 / $12, and the extra rate is applied retroactively across the whole call. The marginal cost of the 200,001st token is not two dollars per million. It is the entire prior prompt, repriced.
That matters more for agents than for anything else, because an agent’s context is not a size you choose once. It grows on its own, turn by turn, as tool results and file reads accumulate. A session that starts comfortably inside the cheap tier can cross the line halfway through and double its per-turn cost for the remainder, without anyone changing a setting.
Anthropic prices the same problem differently, and is explicit about it. Its pricing documentation states that Claude 4.6 and later models “include the full 1M token context window at standard pricing,” adding the parenthetical: “A 900k-token request is billed at the same per-token rate as a 9k-token request.” Claude Opus 5 bills $5 input, $0.50 cache read and $25 output at every context length it supports.
So the two cards converge as context grows. Under 200K, Grok 4.6 cache reads cost $0.50 per million, exactly what Anthropic charges for a cache hit on Opus 5. Over 200K, Grok 4.6 cache reads cost $1.00 per million, double Anthropic’s.
| Context size | Grok 4.6 | Claude Opus 5 |
|---|---|---|
| 50K | $0.04 | $0.09 |
| 100K | $0.07 | $0.13 |
| 150K | $0.11 | $0.18 |
| 200K | $0.28 | $0.23 |
| 250K | $0.34 | $0.28 |
| 300K | $0.41 | $0.32 |
| 350K | $0.47 | $0.37 |
| 400K | $0.54 | $0.42 |
| 450K | $0.60 | $0.47 |
| 500K | $0.67 | $0.51 |
Note what this chart does not say. On a single request made of entirely fresh input, Grok 4.6 stays cheaper above the line as well, because $4 is still less than $5. The crossover exists specifically for cache-heavy workloads, which is to say for agents, which is to say for the workload the model was announced for.
Modeling the agent bill
A per-token rate is not a cost, and this site has argued that at length in why cheaper AI models cost more. The number that reaches a budget is dollars per finished session. So model one.
| Scenario | Cached input | Fresh input | Output | Total |
|---|---|---|---|---|
| Grok 4.5 (cached at $0.30) | $1.35 | $1.00 | $0.45 | $2.80 |
| Grok 4.6 (under 200K) | $2.25 | $1.00 | $0.45 | $3.70 |
| Grok 4.6 (200K and above) | $4.50 | $2.00 | $0.90 | $7.40 |
| Claude Opus 5 (any context length) | $2.25 | $2.50 | $1.88 | $6.63 |
Three things fall out of that picture.
The same work costs 32 percent more on the newer model. Grok 4.5 runs this session for $2.80; Grok 4.6 runs it for $3.70. Nothing about the workload changed and the advertised price did not change either. The cached-input line did.
Grok 4.6 is cheaper than Opus 5, but by 1.8x, not 4.2x. The output rate is 4.2x apart ($6 against $25) and that ratio is what the launch coverage reached for. On this session, output is 12 percent of the Grok bill. The largest line, cached input at 61 percent of the total, is priced identically on both models. A discount that only applies to the small lines is a small discount.
Above 200K, the ranking inverts. The same session costs $7.40 on Grok 4.6 and $6.63 on Claude Opus 5. The model with the cheaper sticker is about 12 percent more expensive, on the workload it was built for, once its context grows past a threshold it will cross on its own.
The capability half did close
None of this is an argument that Grok 4.6 is a weak model. On capability it did what Grok 4.5 could not.
When this site examined Grok 4.5 in July, the verdict was that cheaper-per-finished-task was defensible but Opus-class capability was not, once the provider’s own benchmark harness was stripped out. Grok 4.5 scored 54 on the Artificial Analysis Intelligence Index. Artificial Analysis now measures Grok 4.6 at 61 at high reasoning effort, a seven-point jump in thirty-five days that ties GPT-5.6 Sol at 61 and trails only Claude Fable 5 at 62 and Claude Opus 5 at 63.
The independent cost-to-run figures on the same evaluation are genuinely favorable to xAI. Artificial Analysis spent $1,068.47 running its Intelligence Index on Grok 4.6, against $2,823.25 on GPT-5.6 Sol and $3,836.05 on Claude Opus 5: the same tasks, the same harness, roughly a third of the Opus bill. Artificial Analysis puts Grok 4.6 at $0.84 per average task and notes it “leads on cost efficiency.”
Artificial Analysis also reports that Grok 4.6 resolves long-horizon tasks in about half the turns and a quarter of the input tokens of Claude Opus 5 at max effort, which is the mechanism behind that gap. Token efficiency is real and it compounds with the per-token rate.
Worth flagging one more thing on the benchmark side: xAI’s own launch numbers cite Terminal-Bench v3.0 at 26 percent while some secondary coverage quotes 88.4 percent on Terminal-Bench v2.1. Those are different benchmark versions, not a contradiction, and the site’s note on benchmark saturation covers why version-mismatched comparisons are so easy to make by accident. Always check which revision a score belongs to.
What to do with this
The practical reading is not “avoid Grok 4.6.” It is priced well and it measures well.
The practical reading is that a 500,000-token context window advertised for long-running agents has an economical range of 200,000 tokens, and that the boundary is a cliff rather than a ramp. If you run agents on it:
- Treat 200K as the real ceiling. Compact, summarize or evict context before the prompt crosses the line, the same way you would manage any hard tier boundary. Crossing it costs more than the tokens that crossed it.
- Instrument the cache line, not the headline. Cache reads are the majority of a long-session bill. A dashboard that tracks input and output tokens but not cache reads will not show you the 67 percent rise between 4.5 and 4.6.
- Re-run your own numbers rather than inheriting the ratio. The 4.2x output gap is real and mostly irrelevant to a cache-heavy workload. The AI coding cost calculator and the model registry on this site hold the current verified rates for this comparison.
- Above 200K, Claude Opus 5 is the cheaper cache. Anthropic’s flat rate across a 1M window is a genuine structural advantage for workloads that genuinely need the room, and it is the sort of thing that never shows up in a launch-day price comparison.
The headline held. The bill did not. That distinction is most of what this site is about, and Grok 4.6 is an unusually clean example of it: a model that got meaningfully smarter, kept its advertised price, and quietly became a third more expensive to run for exactly the customers it was announced for.
Frequently asked questions
- How much does Grok 4.6 cost?
- Grok 4.6 costs $2 per million input tokens, $0.50 per million cached input tokens and $6 per million output tokens for prompts under 200,000 tokens. At 200,000 prompt tokens and above, every rate doubles to $4 input, $1 cached and $12 output, applied to all tokens in the request. A faster variant is offered at twice the price.
- Is Grok 4.6 more expensive than Grok 4.5?
- Yes, for cache-heavy workloads. The headline input and output rates are unchanged at $2 and $6 per million, but cached input rose from $0.30 to $0.50 per million, a 67 percent increase. On a modeled 50-turn agent session with a 90 percent cache-hit rate, that makes Grok 4.6 about 32 percent more expensive than Grok 4.5 for identical work.
- What is the Grok 4.6 200K context price cliff?
- Once a prompt reaches 200,000 tokens, xAI bills the entire request at its long-context rates rather than only the tokens above the threshold. A 201,000-token request therefore costs roughly twice a 199,000-token one. Because an agent accumulates context as it works, a session can cross the boundary mid-run and double its per-turn cost without any configuration change.
- Is Grok 4.6 cheaper than Claude Opus 5?
- Under 200,000 tokens of context, yes: a modeled 50-turn agent session costs $3.70 on Grok 4.6 against $6.63 on Claude Opus 5, about 1.8x cheaper rather than the 4.2x the output rates imply. At or above 200,000 tokens the same session costs $7.40 on Grok 4.6, roughly 12 percent more than Claude Opus 5, because Anthropic charges one flat rate across its full 1M-token window while xAI doubles its cached-input rate.
- How smart is Grok 4.6 compared to Claude Opus 5 and GPT-5.6 Sol?
- Artificial Analysis measures Grok 4.6 at 61 on its Intelligence Index at high reasoning effort. That ties GPT-5.6 Sol at 61 and trails Claude Fable 5 at 62 and Claude Opus 5 at 63. It is a seven-point improvement over Grok 4.5, which scored 54.
Sources
xAI (2026). Grok 4.6. SpaceXAI announcement. https://x.ai/news/grok-4-6
xAI (2026). Models and Pricing. SpaceXAI developer documentation. https://docs.x.ai/docs/models
Anthropic (2026). Pricing. Claude developer platform documentation. https://platform.claude.com/docs/en/about-claude/pricing
Anthropic (2026). Models overview. Claude developer platform documentation. https://platform.claude.com/docs/en/about-claude/models/overview
Artificial Analysis (2026). Grok 4.6 (high): Intelligence, Performance and Price Analysis. https://artificialanalysis.ai/models/grok-4-6
Artificial Analysis (2026). Grok 4.6 returns SpaceXAI to the intelligence frontier and leads on cost efficiency. https://artificialanalysis.ai/articles/grok-4-6-benchmarks-and-analysis
Artificial Analysis (2026). Claude Opus 5: Intelligence, Performance and Price Analysis. https://artificialanalysis.ai/models/claude-opus-5
Artificial Analysis (2026). GPT-5.6 Sol (max): Intelligence, Performance and Price Analysis. https://artificialanalysis.ai/models/gpt-5-6-sol
MarkTechPost (2026). SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Coding, and Knowledge Work. https://www.marktechpost.com/2026/08/12/spacexai-releases-grok-4-6/