Capital & Compute

Terminal-Bench 4.0: The $234 Solved Task

Terminal-Bench 4.0 cut to 66 tasks and Opus 5 leads at 51.8 percent. Derived cost per solved task runs from 6 to 234 dollars, a 38x spread.

· ai· benchmarks· evaluation· economics· By Capital & Compute

The Terminal-Bench 4.0 leaderboard publishes something most benchmark boards leave out: the dollar cost of each run. Ten agent-and-model pairs, 330 trials each, and a full-precision total spend next to every score.

Divide that spend by the number of trials the model actually solved and you get the number that decides procurement: the price of one solved task. On the current board it runs from $6.08 to $234.24. The most expensive entry is also tied for last place.

66
Tasks in 4.0
down from 74 in version 3.0
51.82%
Top score
Claude Opus 5, Claude Code, max effort
38x
Cost-per-solved-task spread
$6.08 to $234.24, derived
3 of 10
Entries on the cost frontier
the other seven are dominated

What Terminal-Bench 4.0 actually changed

Terminal-Bench grades an agent on hard command-line work by running a verification script against the machine’s end state. The agent either reached the required state or it did not, so a convincing transcript earns nothing.

Version 4.0, announced on August 28, 2026 by Ryan Marten, is a maintenance release rather than a new benchmark:

  • Resources were calibrated. Every task now carries a flat 8-hour agent timeout, replacing per-task budgets. The announcement reports that frontier models now never or rarely time out.
  • Eight tasks were removed: two for saturation, two for refusals, two because public solutions exist, and two for unresolved quality or platform-compatibility problems. A task counted as saturated when every class in every family of the latest model generation solved it five times out of five.
  • Nineteen tasks were fixed, updating instructions, environments and verifiers.

That leaves 66 tasks, which is 74 minus 8. Two independent checks confirm it: the v4.0.0 release tree on GitHub contains exactly 66 task directories, and every leaderboard row reports 330 trials, which is 66 tasks at five trials each.

The cost of one solved task

Here is the full board, with the derived column on the right.

Model Agent Accuracy ±95% CI Solved Total cost Cost per solved task
Claude Opus 5 Claude Code 51.82% 3.39 171/330 $5,969.11 $34.91
Claude Fable 5 Claude Code 44.55% 3.85 147/330 $7,265.01 $49.42
GLM-5.3 Claude Code 41.82% 3.23 138/330 $2,727.63 $19.77
GPT-5.6 Sol Codex 37.27% 3.78 123/330 $2,541.70 $20.66
Claude Opus 4.8 Claude Code 23.64% 3.56 78/330 $6,481.26 $83.09
GPT-5.6 Terra Codex 21.52% 3.25 71/330 $1,733.52 $24.42
Grok 4.6 Grok Build 20.30% 3.09 67/330 $3,591.58 $53.61
GPT-5.6 Luna Codex 17.27% 2.85 57/330 $346.67 $6.08
Grok 4.5 Grok Build 12.42% 2.62 41/330 $2,094.11 $51.08
Claude Sonnet 5 Claude Code 12.42% 3.06 41/330 $9,603.86 $234.24

The accuracy, confidence interval, solved count and total cost are published by Terminal-Bench. The final column is this site’s arithmetic: published cost divided by published solved trials. It is not a figure tbench.ai reports, and it is not a price anyone quoted.

It is also not a per-task price you would pay in production. A benchmark trial is a worst-case unit of work: a hard, unfamiliar task with an 8-hour budget and no repository context the model has seen before. Read the column as a ratio between models measured under identical conditions, not as a quote.

Terminal-Bench 4.0 cost-efficiency frontierScatter of accuracy against cost per solved task on a log axis. GPT-5.6 Luna sits at 17.27% for $6.08, GLM-5.3 at 41.82% for $19.77 and Claude Opus 5 at 51.82% for $34.91; these three form the frontier. GPT-5.6 Sol, GPT-5.6 Terra, Claude Fable 5, Grok 4.5, Grok 4.6, Claude Opus 4.8 and Claude Sonnet 5 are all dominated.0%10%20%30%40%50%60%$5$10$20$30$50$100$200$300Cost per solved task (log scale)AccuracyGPT-5.6 LunaGLM-5.3GPT-5.6 SolGPT-5.6 TerraOpus 5Fable 5Grok 4.5Grok 4.6Opus 4.8Sonnet 5On the frontierDominated: cheaper and better exists
Terminal-Bench 4.0 cost-efficiency frontier
ModelAccuracyCost per solved taskOn the cost-efficiency frontier
Opus 551.82%$34.91Yes
Fable 544.55%$49.42No
GLM-5.341.82%$19.77Yes
GPT-5.6 Sol37.27%$20.66No
Opus 4.823.64%$83.09No
GPT-5.6 Terra21.52%$24.42No
Grok 4.620.3%$53.61No
GPT-5.6 Luna17.27%$6.08Yes
Grok 4.512.42%$51.08No
Sonnet 512.42%$234.24No
Accuracy against cost per solved task on Terminal-Bench 4.0. Cost is on a log scale because the field spans 38x. The dashed line is the cost-efficiency frontier: only GPT-5.6 Luna, GLM-5.3 and Claude Opus 5 sit on it. Every dimmed point has something on the board that is both cheaper per solved task and more accurate.Source: Derived from the Terminal-Bench 4.0 leaderboard (accuracy, solved count and total cost published; cost per solved task is this site's arithmetic), August 2026

Three entries are non-dominated, meaning nothing else on the board beats them on both axes at once:

  • GPT-5.6 Luna at $6.08 is the floor. It solves only 17.27%, but nothing cheaper exists.
  • GLM-5.3 at $19.77 is the value pick. It reaches 41.82%, and every cheaper entry scores lower.
  • Claude Opus 5 at $34.91 is the capability pick. It is the most accurate score on the board, and every cheaper entry scores lower.

The other seven entries are strictly dominated. GPT-5.6 Sol at $20.66 and 37.27% is beaten on both counts by GLM-5.3, which is cheaper per solved task and 4.5 points more accurate. Claude Fable 5 costs $49.42 to Opus 5’s $34.91 and scores 7.3 points lower. Claude Opus 4.8 is the clearest generational case: it costs 2.4 times what Opus 5 costs per solved task and solves 93 fewer trials.

Sonnet 5 is the most expensive way to fail

The single most striking row is Claude Sonnet 5. It posted the largest absolute bill on the board, $9,603.86, which is $3,634 more than Opus 5 spent. It solved 41 of 330 trials to Opus 5’s 171. Per solved task it costs $234.24, which is 6.7 times Opus 5 and 38.5 times GPT-5.6 Luna.

The cheaper model produced the larger invoice, and a quarter of the results.

The mechanism is in the announcement, and it is token consumption rather than the per-token rate. Terminal-Bench reports Sonnet 5 burning 21.6 billion tokens on its leaderboard run against 6.5 billion for Opus 5, hitting both timeouts and output-token-exceeded errors, and struggling to stay inside its maximum output length. The operators note they did not enable the 128k max-output-tokens setting, which might have helped, and that they saw large variance in how long Sonnet 5 took.

This is the failure mode covered in why cheaper AI models cost more, and it is the reason a per-million-token rate card cannot rank models on its own. A weaker model on a long-horizon task does not fail quickly and cheaply. It retries, re-reads, and re-plans until the budget runs out, and the meter runs the entire time. Sonnet 5 is the cheapest model per token in the Anthropic column of this board and the most expensive per unit of finished work by a factor of six.

The same logic cuts the other way for GPT-5.6 Luna, which consumed 11.6 billion tokens, more than Opus 5, and still billed only $346.67. Cheap tokens plus heavy caching keep a high-volume, low-accuracy run affordable. Volume alone does not predict the bill, and price alone does not predict the bill. Only the product of the two does.

The same name, three different numbers

Terminal-Bench has published three headline top scores in four months, and treating them as a trend line would be a mistake.

Terminal-Bench top score by versionVersion 2.1 in May 2026 had 89 tasks and a top score of 83.4 percent set by GPT-5.5 on Codex. Version 3.0 in July 2026 had 74 tasks and a top score of 34.4 percent set by GPT-5.6 Sol on Codex. Version 4.0 in August 2026 has 66 tasks and a top score of 51.8 percent set by Claude Opus 5 on Claude Code.0%25%50%75%100%different task setdifferent task set83.4%34.4%51.8%v2.1May 202689 tasksGPT-5.5 / Codexv3.0Jul 202674 tasksGPT-5.6 Sol / Codexv4.0Aug 202666 tasksOpus 5 / Claude Code
Terminal-Bench top score by version
VersionReleasedTasksTop scoreSet by
v2.1May 20268983.4%GPT-5.5 / Codex
v3.0Jul 20267434.4%GPT-5.6 Sol / Codex
v4.0Aug 20266651.8%Opus 5 / Claude Code
Terminal-Bench top scores by version. The columns are deliberately not connected: 2.1, 3.0 and 4.0 measure different task sets, so the sequence is three separate measurements that share a name, not a capability trajectory.Source: Terminal-Bench version announcements and the 4.0 leaderboard, May to August 2026

Version 3.0, released on July 30, 2026, was not an edit of 2.1. It was a fresh set of 74 tasks across seven domains, built explicitly because the older tasks had saturated and the leaderboard had compressed into a range too narrow to separate models. The top score fell from about 83% to 34.4% because the work got harder, not because the models got worse. The same team took the same approach with Harbor-Index, where no agent it tested cleared 30%. Between 3.0 and 4.0 the score then rose to 51.82%, partly because a stronger model arrived and partly because the task set lost eight of its members and every task gained an 8-hour budget.

None of this is a defect. It is the maintenance an execution-graded benchmark needs, and Terminal-Bench documents each change in public. But it does mean a launch post citing “Terminal-Bench” with no version attached is quoting a number with no fixed meaning, which is the argument made at length in are AI benchmarks reliable and AI agent benchmarks in 2026. Version 2.1 and version 4.0 differ by more than 30 points on the same name.

How to read the gaps honestly

Two published caveats should change how you rank this board.

The confidence intervals are wide. Every row carries a 95% confidence half-width between 2.62 and 3.85 points. Grok 4.5 and Claude Sonnet 5 both sit at 12.42%, so they are tied. Grok 4.6 at 20.30% and GPT-5.6 Terra at 21.52% are 1.2 points apart on intervals of about 3.2, so that ordering is noise.

Infrastructure alone moves scores by more than some gaps. The resource calibration in 4.0 follows Quantifying infrastructure noise in agentic coding evals, published by Gian Segato at Anthropic in February 2026. Running Terminal-Bench 2.0 across six resource configurations, from strictly enforced to entirely uncapped, produced a 6 percentage point gap between the most- and least-resourced setups. Its recommendation is explicit: leaderboard differences below 3 percentage points deserve skepticism until the evaluation configuration is documented and matched.

Applied here, the safe read is three tiers rather than ten ranks. Opus 5 leads. Fable 5 and GLM-5.3 form a contested second group. Everything from Opus 4.8 down is separated more by confidence interval than by demonstrated capability.

What 4.1 and 5.0 will change

Neither has shipped. Both exist as open milestones on the Terminal-Bench repository, with no due date and no closed issues as of August 29, 2026:

  • Terminal-Bench 4.1 is scoped to verifier improvements, specifically tamper-resistant verifiers, plus task fixes at minor and patch level. Because verifier-only changes are a minor version under the semantic policy, saved artifacts get re-graded rather than re-run, so scores can move without anyone paying for new rollouts. It carries 66 open issues.
  • Terminal-Bench 5.0 is scoped to new tasks alongside significant task modifications to timeouts, resources and the agent environment. That is a breaking change by definition, so it will require full re-runs and will reset the board again. It carries 79 open issues.

Treat any figure attributed to 5.0 before that milestone closes as unverified. The practical consequence for anyone citing these numbers is that the 51.82% ceiling is a snapshot with a known expiry, not a stable fact about Opus 5.

The practical read

If you are choosing a coding agent stack on this evidence, the board supports three conclusions and no more.

  1. Rank on cost per solved task, not on the score column. The accuracy ranking and the efficiency ranking disagree sharply: GLM-5.3 is third on accuracy and second-cheapest per solved task, while Fable 5 is second on accuracy and fourth-most-expensive.
  2. Do not assume the cheap model is the cheap option. Sonnet 5 is the counter-example, and its mechanism, token burn on long-horizon tasks, is a property of the work rather than of the price list. Our cost-per-task tracker models the same effect across harnesses.
  3. Record the version with the number. A Terminal-Bench figure without one is unusable, and 5.0 will invalidate the current board when it lands.

Frequently asked questions

What is Terminal-Bench 4.0?
Terminal-Bench 4.0 is the August 28, 2026 release of the Terminal-Bench agentic command-line benchmark. It contains 66 containerized tasks, each attempted five times for 330 trials per leaderboard entry, with a flat 8-hour agent timeout. It removed 8 tasks and fixed 19 relative to version 3.0.
Which model leads Terminal-Bench 4.0?
Claude Opus 5 running on Claude Code at max effort leads with 51.82%, solving 171 of 330 trials. Claude Fable 5 follows at 44.55% and GLM-5.3 at 41.82%. The 95% confidence half-widths run from 2.62 to 3.85 points, so closely spaced entries are not reliably separated.
How much does it cost to solve one Terminal-Bench task?
Dividing each entry published total cost by its solved trials gives a range from $6.08 for GPT-5.6 Luna to $234.24 for Claude Sonnet 5, a 38x spread. Claude Opus 5 costs $34.91 and GLM-5.3 costs $19.77. These are derived figures, not prices Terminal-Bench publishes, and a benchmark trial is a harder unit of work than a typical production task.
Why is Claude Sonnet 5 the most expensive model on the board?
Because of token volume, not per-token price. Terminal-Bench reports Sonnet 5 consuming 21.6 billion tokens on its run against 6.5 billion for Opus 5, while hitting timeouts and output-token-exceeded errors. Its total bill of $9,603.86 is the largest on the board even though it tied for last place at 12.42%.
Can I compare a Terminal-Bench 2.1 score to a 4.0 score?
No. Version 3.0 replaced the task set entirely with 74 new tasks, and 4.0 cut that to 66 and changed the resource limits. Top scores went from about 83% on 2.1 to 34.4% on 3.0 to 51.82% on 4.0. Those are three separate measurements that share a name.
Has Terminal-Bench 5.0 been released?
Not as of August 29, 2026. Terminal-Bench 5.0 exists as an open GitHub milestone scoped to new tasks and significant environment changes, with 79 open issues, no closed issues and no due date. Version 4.1, scoped to tamper-resistant verifiers, is also open and unshipped.

Sources

  • Marten, R. (2026). Terminal-Bench 4.0: Calibrating task resources, fixing tasks, and removing saturated tasks [announcement]. Terminal-Bench. Verified 2026-08-29. tbench.ai
  • Terminal-Bench (2026). Terminal-Bench 4.0 leaderboard [accuracy, confidence intervals, solved counts, total cost and token counts for all ten entries]. Verified 2026-08-29. tbench.ai
  • Terminal-Bench (2026). Terminal-Bench 3.0 [announcement; 74 tasks across 7 domains, per-version pass rates]. Verified 2026-08-29. tbench.ai
  • Terminal-Bench (2026). Continuous Benchmarks [semantic versioning policy for patch, minor and major releases]. Verified 2026-08-29. tbench.ai
  • Harbor Framework (2026). terminal-bench v4.0.0 [release notes naming the 8 removed tasks; task tree contains 66 directories]. GitHub. Verified 2026-08-29. github.com
  • Segato, G. (2026). Quantifying infrastructure noise in agentic coding evals [engineering research; 6 percentage-point gap across six resource configurations on Terminal-Bench 2.0]. Anthropic, February 5, 2026. Verified 2026-08-29. anthropic.com
  • Merrill, M. A., et al. (2026). Terminal-Bench: Benchmarking agents on hard, realistic tasks in command line interfaces [preprint]. arXiv:2601.11868. arxiv.org

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← Back to Models & benchmarks