Capital & Compute

Claude Opus 5.5: Pricing, Benchmarks, Cost

Claude Opus 5.5 lists $4 and $20 per million tokens, 20 percent under Opus 5. Anthropic claims 40 percent cheaper to run. Here is the missing half.

· Updated September 29, 2026· ai· pricing· benchmarks· economics· By Capital & Compute
Bar chart of Claude Opus 5.5 gains over Opus 5 on eight benchmarks, from 1.06x to 2.02x.

Anthropic released Claude Opus 5.5 on September 22, 2026 at $4 per million input tokens and $20 per million output, against the $5 and $25 that Opus 5 charged. Every line of the rate card fell 20%, and the cache read fell 60%, from $0.50 to $0.20. In sixty-seven tracked model releases this is the first Anthropic flagship to launch under the price of the model it replaced.

The company also says it costs “40% less to run than Opus 5 on typical workloads.” That number is larger than the rate card can produce. The rate cut accounts for about 24 points of it on Anthropic’s own blended basis. The other 16 depend entirely on a claim about token efficiency that nobody outside Anthropic has measured yet, and the one independent reading available points in an awkward direction.

Opus 5 to Opus 5.5, as measured by eight vendor scoreboardsEight scoreboards measuring the same generational step, each shown as a multiple of the Opus 5 result. Terminal-Bench-Science 0.1 moves from 29.0 to 58.7, a 2.02x step. AutomationBench moves 26.9 to 40.0, 1.49x. Terminal-Bench 4.0 moves 52.3 to 66.4, 1.27x. CursorBench 4.0 moves 46.6 to 57.8, 1.24x. FrontierCode v1.1 moves 48.0 to 54.4, 1.13x. OSWorld 2.0 moves 74.0 to 81.8, 1.11x. Chartography moves 83.4 to 89.0, 1.07x. Humanity's Last Exam moves 63.6 to 67.7, 1.06x. The same release therefore reads as a 6 percent improvement or a doubling depending on which board is quoted.1x1.25x1.5x2x2.5xMultiple of the Opus 5 result, log scale. 1x is Opus 5.Terminal-Bench-Science 0.129.0 to 58.7 percent2.02xAutomationBench26.9 to 40.0 percent1.49xTerminal-Bench 4.052.3 to 66.4 percent1.27xCursorBench 4.046.6 to 57.8 percent1.24xFrontierCode v1.148.0 to 54.4 percent1.13xOSWorld 2.074.0 to 81.8 percent1.11xChartography83.4 to 89.0 percent, with tools1.07xHumanity's Last Exam63.6 to 67.7 percent, with tools1.06x
Opus 5 to Opus 5.5, as measured by eight vendor scoreboards
ScoreboardPrevious modelNew modelMultiple
Terminal-Bench-Science 0.1 (29.0 to 58.7 percent)2958.72.02x
AutomationBench (26.9 to 40.0 percent)26.9401.49x
Terminal-Bench 4.0 (52.3 to 66.4 percent)52.366.41.27x
CursorBench 4.0 (46.6 to 57.8 percent)46.657.81.24x
FrontierCode v1.1 (48.0 to 54.4 percent)4854.41.13x
OSWorld 2.0 (74.0 to 81.8 percent)7481.81.11x
Chartography (83.4 to 89.0 percent, with tools)83.4891.07x
Humanity's Last Exam (63.6 to 67.7 percent, with tools)63.667.71.06x
One generational step, measured eight different ways. Each bar runs from the Opus 5 result out to the Opus 5.5 result on the same board, on a log axis because the boards disagree by nearly 2x about how large the step was. All figures are Anthropic reporting on Anthropic.Source: Anthropic, Claude Opus 5.5 announcement (vendor-reported), read September 23, 2026

Pick your headline. The same release is a 6% improvement on Humanity’s Last Exam and a doubling on Terminal-Bench-Science. Anthropic quotes all of them, which is fair; the spread is the reason a single benchmark number should never carry a purchasing decision.

How much does Claude Opus 5.5 cost?

Every rate below is from Anthropic’s pricing documentation, read September 23, 2026.

Billing line Opus 5.5 Opus 5 Change
Input per Mtok $4 $5 -20%
Output per Mtok $20 $25 -20%
Cache read per Mtok $0.20 $0.50 -60%
Cache write, 5 minute $5 $6.25 -20%
Cache write, 1 hour $8 $10 -20%
Batch input / output $2 / $10 $2.50 / $12.50 -20%
Fast mode in / out $8 / $40 $10 / $50 -20%

The cache read is the line that did not simply follow the others. Anthropic bills a cache hit as a multiplier on base input, normally 0.1x. On Opus 5.5 that multiplier itself halves to 0.05x, which is why a 20% cut to input produces a 60% cut to the cache read. Only Claude Fable 5.1 reads cheaper relative to its own input rate, at 0.025x, and at $0.25 per Mtok it is no longer the cheapest cache read on the Claude ladder in absolute dollars.

That matters more than it sounds. In a long agent session most input tokens are cache reads, not fresh ones, which is why the cost-per-task calculator weights them heavily.

The 40% claim, and what would have to be true

Anthropic’s own framing is that Opus 5.5 “costs 40% less to run than Opus 5 on typical workloads.” Two separate things can produce that: a lower rate, and fewer tokens spent finishing the same work. Only the first is verifiable today.

Artificial Analysis prices models on a blend of 7 parts cache hit, 2 parts fresh input and 1 part output. Run the two rate cards through it:

  • Opus 5: (0.7 x $0.50) + (0.2 x $5) + (0.1 x $25) = $3.85 per million tokens
  • Opus 5.5: (0.7 x $0.20) + (0.2 x $4) + (0.1 x $20) = $2.94 per million tokens

The second figure is exactly what Artificial Analysis publishes for Opus 5.5, which is the check that the blend is being applied the way they apply it. The cut is 23.6%.

To reach 40%, the effective cost of a task has to land at 60% of $3.85, or $2.31. At a blended $2.94, that requires the same work to finish in 78.6% of the tokens: a 21% reduction in tokens per task. That is the entire unverified half of the claim, stated as a number you could go and test.

Terminal-Bench has not run it

Anthropic reports 66.4% on Terminal-Bench 4.0, which would top the public board by more than eight points. A read of the Terminal-Bench 4.0 leaderboard on September 23, 2026 returned fifteen model rows and no Opus 5.5 entry. The board currently leads with GPT-6 Astra at 58.2% and Claude Fable 5.1 at 57.9%.

This is not a small gap in the record. The Terminal-Bench board publishes each entry’s total run cost alongside its solved count, which is what makes a cost per solved task derivable at all. On that board Opus 5 solves a task for $34.91 and GPT-6 Astra for $17.02, a comparison covered in full in the $234 solved task. Until Terminal-Bench runs Opus 5.5 there is no independent accuracy to pair with a published cost, so no equivalent figure exists for it, and the 40% claim stays untested on the one public benchmark that measures cost directly.

What the effort dial actually buys

Opus 5.5 exposes five effort settings. Artificial Analysis scores each separately and publishes what each costs to run its Intelligence Index, which turns the dial into a price list.

Effort Intelligence Index Cost per index task Points above low Cost above low
low 42 $0.55 0 1.0x
medium 51 $1.34 +9 2.4x
high 54 $1.82 +12 3.3x
xhigh 56 $3.46 +14 6.3x
max 58 $5.98 +16 10.9x
Claude Opus 5.5 effort settings: cost against index scoreCost to run the Artificial Analysis Intelligence Index at each Claude Opus 5.5 effort setting, on a log axis. Low effort costs $0.55 per index task and scores 42. Medium costs $1.34 and scores 51. High costs $1.82 and scores 54. Extended-high costs $3.46 and scores 56. Max costs $5.98 and scores 58. The full range is 10.9 times, and the step from high to max costs 3.3 times as much for four additional index points.$0.5$1$2$5lowIntelligence Index 42$0.55mediumIntelligence Index 51$1.34highIntelligence Index 54$1.82xhighIntelligence Index 56$3.46maxIntelligence Index 58$5.98
Claude Opus 5.5 effort settings: cost against index score
Effort settingCost per Artificial Analysis index task
low (Intelligence Index 42)$0.55
medium (Intelligence Index 51)$1.34
high (Intelligence Index 54)$1.82
xhigh (Intelligence Index 56)$3.46
max (Intelligence Index 58)$5.98
What each effort setting costs to run the Artificial Analysis Intelligence Index, on a log axis because the dial spans 10.9x. The index score each setting buys is under its label. The last four points cost more than the first twelve.Source: Artificial Analysis, read September 23, 2026

The dial is steeply non-linear and the knee is obvious. Going from low to high buys 12 of the 16 available index points for 3.3 times the cost. Going from high to max buys the last 4 points for another 3.3 times on top. If a workload is not failing at high effort, max is paying triple for a four-point aggregate difference.

At max effort, Opus 5.5 scores 58 and ranks first of 212 models on that index. That is a genuine result and worth stating plainly.

Where Opus 5.5 does not win

A launch post that finds nothing is a press release. On Anthropic’s own numbers there are two clear losses, both to OpenAI:

  • AutomationBench: Opus 5.5 scores 40.0% against GPT-6 Astra at 41.4%.
  • Terminal-Bench-Science 0.1: Opus 5.5 scores 58.7% against GPT-6 Astra at 64.6%, a six-point deficit on the very board where Opus 5.5 posts its largest generational gain.

There is also the verbosity reading. Artificial Analysis records that Opus 5.5 generates 260 million output tokens to complete its Intelligence Index, against a median of 88 million across the models it tracks, and describes the model as very verbose. This is not proof that the token-efficiency claim is wrong: Opus 5 was itself unusually verbose, burning 6.53 billion tokens on Terminal-Bench 4.0 where GPT-6 Astra burned 1.53 billion, so Opus 5.5 could be a large improvement on its predecessor and still sit well above the median. No published side-by-side token count exists. That is precisely the problem.

Should you switch?

For anyone already on Opus 5, the answer is unusually simple, and it does not depend on any contested claim. Opus 5.5 is cheaper on every billing line, scores higher on every board Anthropic published, and runs on the same 1M-token context. There is no configuration in which Opus 5 is the better buy. The same release also set off a wave of user-made motion graphics, and the tools those clips are rendered in are ranked here.

The harder questions are the ones the rate cut reopens:

  • Coming from Sonnet 5 at $2 and $10, the escalation premium is now 2x rather than the 2.5x Opus 5 represented. Claude Sonnet 5.5, released September 28 on the same rate card, narrows the capability gap to two index points but costs more than Opus 5.5 per Artificial Analysis index task at max effort. That is a materially different trade, and it is worth re-running against your own task mix in the cost-per-task model rather than reusing a conclusion reached at the old ratio. The full routing comparison, including why the task-cost gap is closer to 24% than 2x, is in Claude Opus 5.5 vs Sonnet 5.
  • Coming from GPT-6 Astra at $10 and $50, Opus 5.5 is 60% cheaper per token, but Astra currently holds the independent Terminal-Bench cost-per-solved-task advantage by a wide margin. See the GPT-6 Astra pricing breakdown for why the cheaper sticker has repeatedly lost that comparison.
  • Coming from the OpenAI reasoning tier, the parity lasted a day. Opus 5.5 matched GPT-5.6 Sol at $4 and $20, but OpenAI replaced that tier hours later with GPT-6 Sol at $2 and $10, half of Opus 5.5 on both lines and level with Claude Sonnet 5. Opus 5.5 scores higher on the Artificial Analysis index, 58 at max against 48, so this is now a capability-for-price trade rather than a tie.

Current rates for every model tracked here, each dated and sourced, sit on the AI model tracker.

Frequently asked questions

How much does Claude Opus 5.5 cost?
Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, with a $0.20 cache read and a $5 cache write at five minutes. That is 20 percent under Claude Opus 5 on every line except the cache read, which is 60 percent lower. Batch pricing halves it to $2 and $10, and fast mode doubles it to $8 and $40.
Is Claude Opus 5.5 cheaper than Opus 5?
Yes, on every billing line. On the blended basis Artificial Analysis uses, 7 parts cache hit to 2 parts input to 1 part output, Opus 5.5 works out at $2.94 per million tokens against $3.85 for Opus 5, a 23.6 percent reduction. Anthropic claims 40 percent less to run, which requires the model to also finish work in about 21 percent fewer tokens.
What did Claude Opus 5.5 score on Terminal-Bench 4.0?
Anthropic reports 66.4 percent, against 55.8 percent for Claude Fable 5.1, 57.9 percent for GPT-6 Astra and 52.3 percent for Claude Opus 5. That figure is vendor-reported. The independent Terminal-Bench 4.0 leaderboard had no Claude Opus 5.5 entry as of September 23, 2026, so it has not been reproduced outside Anthropic.
Where does Claude Opus 5.5 rank on Artificial Analysis?
First of 212 models, at 58 on the Artificial Analysis Intelligence Index running at max effort with default fallback, read September 23, 2026. Claude Fable 5.1 and GPT-6 Astra both score 53 at max effort. Artificial Analysis rescaled this index during 2026, so 58 cannot be compared against index figures quoted earlier in the year.
Which effort level should I use on Claude Opus 5.5?
High effort is the value knee. It scores 54 on the Artificial Analysis Intelligence Index at $1.82 per index task, against 42 at $0.55 for low and 58 at $5.98 for max. Moving from low to high buys 12 of the 16 available points for 3.3 times the cost; moving from high to max buys the last 4 points for another 3.3 times on top.
Should I switch from Claude Opus 5 to Opus 5.5?
Yes. Opus 5.5 is cheaper on every billing line, scores higher on all eight benchmarks Anthropic published, and keeps the same 1M-token context window. Opus 5 remains callable on the API for workloads pinned to a fixed model, but there is no cost or capability case for choosing it.

Sources

Anthropic (2026). Claude Opus 5.5. Provider announcement (vendor-reported benchmarks). https://www.anthropic.com/claude-opus-5-5 Verified 2026-09-23.

Anthropic (2026). Pricing. Claude Platform documentation (input, output, cache, batch and fast-mode rates; cache-read multipliers). https://platform.claude.com/docs/en/about-claude/pricing Verified 2026-09-23.

Artificial Analysis (2026). LLM Leaderboard. Independent benchmarking service (Intelligence Index by effort level, cost per index task, blended price, output-token counts). https://artificialanalysis.ai/leaderboards/models Verified 2026-09-23.

Terminal-Bench (2026). Terminal-Bench 4.0 leaderboard. Independent benchmark (resolution rate, run cost and token counts per entry; no Claude Opus 5.5 entry at time of reading). https://www.tbench.ai/leaderboard/terminal-bench/4.0 Verified 2026-09-23.

The Artificial Analysis and Terminal-Bench pages are client-rendered and return an empty table to a plain HTTP fetch. Both were read in a headless browser on September 23, 2026.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← Back to Models & benchmarks