Capital & Compute

Terminal-Bench 4.0: The $234 Solved Task

Terminal-Bench 4.0 now runs 18 entries and GPT-6 Astra leads at 58.2 percent. Derived cost per solved task runs from 6 to 234 dollars, a 38x spread.

· Updated September 4, 2026· ai· benchmarks· evaluation· economics· By Capital & Compute

The Terminal-Bench 4.0 leaderboard publishes something most benchmark boards leave out: the dollar cost of each run. Eighteen agent-and-model pairs, 330 trials each, and a full-precision total spend next to every score.

Divide that spend by the number of trials the model actually solved and you get the number that decides procurement: the price of one solved task. On the current board it runs from $6.08 to $234.24. The most expensive entry is also tied for last place.

66
Tasks in 4.0
down from 74 in version 3.0
58.18%
Top score
GPT-6 Astra, Codex, max effort
38x
Cost-per-solved-task spread
$6.08 to $234.24, derived
5 of 18
Entries on the cost frontier
four of the five are Astra effort settings

What Terminal-Bench 4.0 actually changed

Terminal-Bench grades an agent on hard command-line work by running a verification script against the machine’s end state. The agent either reached the required state or it did not, so a convincing transcript earns nothing.

Version 4.0, announced on August 28, 2026 by Ryan Marten, is a maintenance release rather than a new benchmark:

  • Resources were calibrated. Every task now carries a flat 8-hour agent timeout, replacing per-task budgets. The announcement reports that frontier models now never or rarely time out.
  • Eight tasks were removed: two for saturation, two for refusals, two because public solutions exist, and two for unresolved quality or platform-compatibility problems. A task counted as saturated when every class in every family of the latest model generation solved it five times out of five.
  • Nineteen tasks were fixed, updating instructions, environments and verifiers.

That leaves 66 tasks, which is 74 minus 8. Two independent checks confirm it: the v4.0.0 release tree on GitHub contains exactly 66 task directories, and every leaderboard row reports 330 trials, which is 66 tasks at five trials each.

The cost of one solved task

Here is the full board, with the derived column on the right.

Model Agent Accuracy ±95% CI Solved Total cost Cost per solved task
GPT-6 Astra (max) Codex 58.18% 2.79 192/330 $3,267.18 $17.02
Claude Fable 5.1 Claude Code 57.88% 3.83 191/330 $6,243.50 $32.69
GPT-6 Astra (xhigh) Codex 57.88% 2.71 191/330 $2,350.51 $12.31
GPT-6 Astra (high) Codex 57.88% 2.98 191/330 $2,269.42 $11.88
GPT-6 Astra (medium) Codex 54.24% 2.74 179/330 $1,914.80 $10.70
Claude Opus 5 Claude Code 51.82% 3.39 171/330 $5,969.11 $34.91
GPT-6 Astra (low) Codex 50.61% 2.83 167/330 $1,557.30 $9.33
Claude Fable 5 Claude Code 44.55% 3.85 147/330 $7,265.01 $49.42
GLM-5.3 Claude Code 41.82% 3.23 138/330 $2,727.63 $19.77
GPT-5.6 Sol Codex 37.27% 3.78 123/330 $2,541.70 $20.66
Claude Opus 4.8 Claude Code 23.64% 3.56 78/330 $6,481.26 $83.09
GPT-5.6 Terra Codex 21.52% 3.25 71/330 $1,733.52 $24.42
Grok 4.6 Grok Build 20.30% 3.09 67/330 $3,591.58 $53.61
Gemini 3.8 Flash mini-SWE-agent 19.09% 3.43 63/330 $1,828.77 $29.03
GPT-5.6 Luna Codex 17.27% 2.85 57/330 $346.67 $6.08
Grok 4.5 Grok Build 12.42% 2.62 41/330 $2,094.11 $51.08
Claude Sonnet 5 Claude Code 12.42% 3.06 41/330 $9,603.86 $234.24
Gemini 3.7 Flash mini-SWE-agent 11.21% 2.44 37/330 $1,261.87 $34.10

The accuracy, confidence interval, solved count and total cost are published by Terminal-Bench. The final column is this site’s arithmetic: published cost divided by published solved trials. It is not a figure tbench.ai reports, and it is not a price anyone quoted.

It is also not a per-task price you would pay in production. A benchmark trial is a worst-case unit of work: a hard, unfamiliar task with an 8-hour budget and no repository context the model has seen before. Read the column as a ratio between models measured under identical conditions, not as a quote.

Terminal-Bench 4.0 cost-efficiency frontierScatter of accuracy against cost per solved task on a log axis for eighteen entries. GPT-5.6 Luna sits at 17.27% for $6.08, GPT-6 Astra at low effort 50.61% for $9.33, at medium 54.24% for $10.70, at high 57.88% for $11.88 and at max 58.18% for $17.02; these five form the frontier. Claude Fable 5.1 at 57.88% for $32.69, Claude Opus 5 at 51.82% for $34.91, GLM-5.3 at 41.82% for $19.77, GPT-5.6 Sol, GPT-5.6 Terra, Gemini 3.8 Flash, Gemini 3.7 Flash, Claude Fable 5, Grok 4.5, Grok 4.6, Claude Opus 4.8, Claude Sonnet 5 and Astra at xhigh are all dominated.0%10%20%30%40%50%60%70%$5$10$20$30$50$100$200$300Cost per solved task (log scale)AccuracyGPT-5.6 LunaAstra (low)Astra (medium)Astra (high)Astra (xhigh)Astra (max)GLM-5.3GPT-5.6 SolGPT-5.6 TerraGemini 3.8 FlashFable 5.1Gemini 3.7 FlashOpus 5Fable 5Grok 4.5Grok 4.6Opus 4.8Sonnet 5On the frontierDominated: cheaper and better exists
Terminal-Bench 4.0 cost-efficiency frontier
ModelAccuracyCost per solved taskOn the cost-efficiency frontier
Astra (max)58.18%$17.02Yes
Fable 5.157.88%$32.69No
Astra (xhigh)57.88%$12.31No
Astra (high)57.88%$11.88Yes
Astra (medium)54.24%$10.70Yes
Opus 551.82%$34.91No
Astra (low)50.61%$9.33Yes
Fable 544.55%$49.42No
GLM-5.341.82%$19.77No
GPT-5.6 Sol37.27%$20.66No
Opus 4.823.64%$83.09No
GPT-5.6 Terra21.52%$24.42No
Grok 4.620.3%$53.61No
Gemini 3.8 Flash19.09%$29.03No
GPT-5.6 Luna17.27%$6.08Yes
Grok 4.512.42%$51.08No
Sonnet 512.42%$234.24No
Gemini 3.7 Flash11.21%$34.10No
Accuracy against cost per solved task on Terminal-Bench 4.0. Cost is on a log scale because the field spans 38x. The dashed line is the cost-efficiency frontier: GPT-5.6 Luna sets the floor and the four cheapest GPT-6 Astra effort settings take the rest of it. Every dimmed point has something on the board that is both cheaper per solved task and more accurate.Source: Derived from the Terminal-Bench 4.0 leaderboard (accuracy, solved count and total cost published; cost per solved task is this site's arithmetic), re-read September 4, 2026

Five entries are non-dominated, meaning nothing else on the board beats them on both axes at once. Four of the five are the same model at different settings:

  • GPT-5.6 Luna at $6.08 is still the floor. It solves only 17.27%, but nothing cheaper exists.
  • GPT-6 Astra at low effort, $9.33 solves 50.61%, which is more than the model that led this board a week ago, for a quarter of the cost per solved task.
  • GPT-6 Astra at medium, $10.70 reaches 54.24%.
  • GPT-6 Astra at high, $11.88 reaches 57.88%, tying Claude Fable 5.1 at 36% of the cost.
  • GPT-6 Astra at max, $17.02 is the capability pick and the most accurate entry on the board.

The arrival of one model knocked two entries off the frontier. GLM-5.3 at $19.77 and 41.82% was the value pick in August; Astra at medium effort is now both cheaper and 12 points more accurate. Claude Opus 5 at $34.91 was the capability pick; Astra at low effort costs a quarter as much and scores slightly higher, and Astra at max is more accurate still.

The other thirteen entries are strictly dominated. Astra at xhigh is dominated by Astra at high, which posts the identical 57.88% for $0.43 less per solved task. Claude Fable 5.1 matches that same accuracy and costs $32.69, though it does so inside a different agent harness. Claude Opus 4.8 remains the clearest generational case: it costs 2.4 times what Opus 5 costs per solved task and solves 93 fewer trials.

Sonnet 5 is the most expensive way to fail

The single most striking row is Claude Sonnet 5. It posted the largest absolute bill on the board, $9,603.86, which is $3,634 more than Opus 5 spent. It solved 41 of 330 trials to Opus 5’s 171. Per solved task it costs $234.24, which is 6.7 times Opus 5 and 38.5 times GPT-5.6 Luna.

The cheaper model produced the larger invoice, and a quarter of the results.

The mechanism is in the announcement, and it is token consumption rather than the per-token rate. Terminal-Bench reports Sonnet 5 burning 21.6 billion tokens on its leaderboard run against 6.5 billion for Opus 5, hitting both timeouts and output-token-exceeded errors, and struggling to stay inside its maximum output length. The operators note they did not enable the 128k max-output-tokens setting, which might have helped, and that they saw large variance in how long Sonnet 5 took.

This is the failure mode covered in why cheaper AI models cost more, and it is the reason a per-million-token rate card cannot rank models on its own. A weaker model on a long-horizon task does not fail quickly and cheaply. It retries, re-reads, and re-plans until the budget runs out, and the meter runs the entire time. Sonnet 5 is the cheapest model per token in the Anthropic column of this board and the most expensive per unit of finished work by a factor of six.

The same logic cuts the other way for GPT-5.6 Luna, which consumed 11.6 billion tokens, more than Opus 5, and still billed only $346.67. Cheap tokens plus heavy caching keep a high-volume, low-accuracy run affordable. Volume alone does not predict the bill, and price alone does not predict the bill. Only the product of the two does.

GPT-6 Astra is the same principle read from the opposite end, and it is the reason the board reordered. Astra is the joint most expensive model here per token, at $10 input and $50 output, matching Claude Fable 5.1 exactly. It consumed 1.53 billion tokens at max effort, the smallest total on the board by a wide margin: Fable 5.1 used 2.75 billion, Opus 5 used 6.53 billion and Sonnet 5 used 21.6 billion. Per solved task that works out to 8.0 million tokens for Astra against 14.4 million for Fable 5.1, 38.2 million for Opus 5 and 525.9 million for Sonnet 5. A model can charge 2.5 times more per token than its own predecessor and still be cheaper to finish the work with, which is exactly what happened between GPT-5.6 Sol and Astra.

The same name, three different numbers

Terminal-Bench has published three headline top scores in four months, and treating them as a trend line would be a mistake.

Terminal-Bench top score by versionVersion 2.1 in May 2026 had 89 tasks and a top score of 83.4 percent set by GPT-5.5 on Codex. Version 3.0 in July 2026 had 74 tasks and a top score of 34.4 percent set by GPT-5.6 Sol on Codex. Version 4.0 has 66 tasks and a top score of 58.2 percent set by GPT-6 Astra on Codex in September 2026.0%25%50%75%100%different task setdifferent task set83.4%34.4%58.2%v2.1May 202689 tasksGPT-5.5 / Codexv3.0Jul 202674 tasksGPT-5.6 Sol / Codexv4.0Aug 202666 tasksGPT-6 Astra / Codex
Terminal-Bench top score by version
VersionReleasedTasksTop scoreSet by
v2.1May 20268983.4%GPT-5.5 / Codex
v3.0Jul 20267434.4%GPT-5.6 Sol / Codex
v4.0Aug 20266658.2%GPT-6 Astra / Codex
Terminal-Bench top scores by version. The columns are deliberately not connected: 2.1, 3.0 and 4.0 measure different task sets, so the sequence is three separate measurements that share a name, not a capability trajectory. The 4.0 figure is the September 2026 board, led by GPT-6 Astra.Source: Terminal-Bench version announcements and the 4.0 leaderboard, May to September 2026

Version 3.0, released on July 30, 2026, was not an edit of 2.1. It was a fresh set of 74 tasks across seven domains, built explicitly because the older tasks had saturated and the leaderboard had compressed into a range too narrow to separate models. The top score fell from about 83% to 34.4% because the work got harder, not because the models got worse. The same team took the same approach with Harbor-Index, where no agent it tested cleared 30%. Between 3.0 and 4.0 the score then rose to 51.82%, partly because a stronger model arrived and partly because the task set lost eight of its members and every task gained an 8-hour budget. Within 4.0 it has since risen again to 58.18%, and that movement is clean: the task set did not change between August 29 and September 4, only the models running against it did.

None of this is a defect. It is the maintenance an execution-graded benchmark needs, and Terminal-Bench documents each change in public. But it does mean a launch post citing “Terminal-Bench” with no version attached is quoting a number with no fixed meaning, which is the argument made at length in are AI benchmarks reliable and AI agent benchmarks in 2026. Version 2.1 and version 4.0 differ by more than 30 points on the same name.

How to read the gaps honestly

Two published caveats should change how you rank this board.

The confidence intervals are wide. Every row carries a 95% confidence half-width between 2.44 and 3.85 points. That matters more now than it did in August, because the top of the board is inside one. GPT-6 Astra at max leads Claude Fable 5.1 by 0.30 points on an interval of 2.79, and leads its own high-effort run by the same 0.30 on an interval of 2.98. Those three are a tie, not a ranking. Lower down, Grok 4.5 and Claude Sonnet 5 both sit at 12.42%, and Grok 4.6 at 20.30% against GPT-5.6 Terra at 21.52% is 1.2 points on intervals of about 3.2, so that ordering is noise too.

Infrastructure alone moves scores by more than some gaps. The resource calibration in 4.0 follows Quantifying infrastructure noise in agentic coding evals, published by Gian Segato at Anthropic in February 2026. Running Terminal-Bench 2.0 across six resource configurations, from strictly enforced to entirely uncapped, produced a 6 percentage point gap between the most- and least-resourced setups. Its recommendation is explicit: leaderboard differences below 3 percentage points deserve skepticism until the evaluation configuration is documented and matched.

Applied here, the safe read is three tiers rather than eighteen ranks. GPT-6 Astra at high effort or above and Claude Fable 5.1 lead, and are not separable from each other. Astra at medium, Claude Opus 5 and Astra at low form a contested second group between 50% and 55%. Everything from Claude Fable 5 down is separated more by confidence interval and by harness than by demonstrated capability. The harness point is worth holding onto: Astra ran in Codex, the Claude models in Claude Code, and the Gemini models in mini-SWE-agent, so a cross-column comparison carries the agent as well as the model.

What 4.1 and 5.0 will change

Neither has shipped. Both exist as open milestones on the Terminal-Bench repository, with no due date and no closed issues as of August 29, 2026:

  • Terminal-Bench 4.1 is scoped to verifier improvements, specifically tamper-resistant verifiers, plus task fixes at minor and patch level. Because verifier-only changes are a minor version under the semantic policy, saved artifacts get re-graded rather than re-run, so scores can move without anyone paying for new rollouts. It carries 66 open issues.
  • Terminal-Bench 5.0 is scoped to new tasks alongside significant task modifications to timeouts, resources and the agent environment. That is a breaking change by definition, so it will require full re-runs and will reset the board again. It carries 79 open issues.

Treat any figure attributed to 5.0 before that milestone closes as unverified. The practical consequence for anyone citing these numbers is that the 58.18% ceiling is a snapshot with a known expiry, not a stable fact about GPT-6 Astra. It replaced a 51.82% ceiling that had stood for six days.

The practical read

If you are choosing a coding agent stack on this evidence, the board supports three conclusions and no more.

  1. Rank on cost per solved task, not on the score column. The accuracy ranking and the efficiency ranking disagree sharply. Claude Fable 5.1 is joint-second on accuracy and eleventh-cheapest per solved task; GPT-6 Astra at high effort matches its score for a third of the money. If you rank on the score column alone you will pay roughly three times more than you need to for the same result.
  2. Do not assume the cheap model is the cheap option, or that the expensive one is expensive. Sonnet 5 is the first counter-example: cheapest Anthropic model per token, most expensive per unit of finished work by a factor of six. GPT-6 Astra is the second, running the other way: joint most expensive per token and the cheapest capable entry on the board. Both are explained by token burn on long-horizon tasks, which is a property of the work rather than of the price list. The cost-per-task tracker models the same effect across harnesses.
  3. Record the version, the effort and the harness with the number. A Terminal-Bench figure without a version is unusable, and 5.0 will invalidate the current board when it lands. That versioning discipline is why Terminal-Bench sits at the top of the benchmark trust ranking. Astra alone occupies five rows spanning 7.6 accuracy points and an 82% range in cost per solved task, so “GPT-6 Astra on Terminal-Bench 4.0” is not yet a specific enough citation to be worth anything.

Frequently asked questions

What is Terminal-Bench 4.0?
Terminal-Bench 4.0 is the August 28, 2026 release of the Terminal-Bench agentic command-line benchmark. It contains 66 containerized tasks, each attempted five times for 330 trials per leaderboard entry, with a flat 8-hour agent timeout. It removed 8 tasks and fixed 19 relative to version 3.0.
Which model leads Terminal-Bench 4.0?
As of September 4, 2026, GPT-6 Astra running on Codex at max effort leads with 58.18%, solving 192 of 330 trials. Claude Fable 5.1 and GPT-6 Astra at xhigh and high effort all sit at 57.88%, and Claude Opus 5 is sixth at 51.82%. The 95% confidence half-widths run from 2.44 to 3.85 points, so the top four entries are a statistical tie rather than a ranking. Astra ran in Codex and the Claude models in Claude Code, so the comparison carries the agent harness as well as the model.
How much does it cost to solve one Terminal-Bench task?
Dividing each entry published total cost by its solved trials gives a range from $6.08 for GPT-5.6 Luna to $234.24 for Claude Sonnet 5, a 38x spread. Among capable entries, GPT-6 Astra costs $11.88 per solved task at high effort and $17.02 at max, Claude Fable 5.1 costs $32.69 and Claude Opus 5 costs $34.91. These are derived figures, not prices Terminal-Bench publishes, and a benchmark trial is a harder unit of work than a typical production task.
Why is Claude Sonnet 5 the most expensive model on the board?
Because of token volume, not per-token price. Terminal-Bench reports Sonnet 5 consuming 21.6 billion tokens on its run against 6.5 billion for Opus 5 and 1.53 billion for GPT-6 Astra, while hitting timeouts and output-token-exceeded errors. Its total bill of $9,603.86 is the largest on the board even though it tied for last place at 12.42%.
Can I compare a Terminal-Bench 2.1 score to a 4.0 score?
No. Version 3.0 replaced the task set entirely with 74 new tasks, and 4.0 cut that to 66 and changed the resource limits. Top scores went from about 83% on 2.1 to 34.4% on 3.0 to 58.18% on 4.0. Those are three separate measurements that share a name. Movement within a single version is comparable: the 4.0 top score rose from 51.82% to 58.18% between August 29 and September 4 with no change to the task set, only new models entering.
How did GPT-6 Astra change the Terminal-Bench 4.0 board?
OpenAI released GPT-6 Astra on September 3, 2026 and Terminal-Bench posted five runs the same day, one per reasoning-effort level. Astra at max took the top score at 58.18%, and its low, medium, high and max settings took four of the five places on the cost-efficiency frontier, displacing both Claude Opus 5 and GLM-5.3. The board grew from ten entries to eighteen. Astra costs the most per token of any model on the board and the least per solved task of any capable one, because it used 1.53 billion tokens where Opus 5 used 6.53 billion.
Has Terminal-Bench 5.0 been released?
Not as of August 29, 2026. Terminal-Bench 5.0 exists as an open GitHub milestone scoped to new tasks and significant environment changes, with 79 open issues, no closed issues and no due date. Version 4.1, scoped to tamper-resistant verifiers, is also open and unshipped.

Sources

  • Marten, R. (2026). Terminal-Bench 4.0: Calibrating task resources, fixing tasks, and removing saturated tasks [announcement]. Terminal-Bench. Verified 2026-08-29. tbench.ai
  • Terminal-Bench (2026). Terminal-Bench 4.0 leaderboard [accuracy, confidence intervals, solved counts, total cost and token counts for all eighteen entries; leaderboard timestamp 2026-09-03 21:34 UTC]. Verified 2026-09-04. tbench.ai
  • OpenAI (2026). API pricing [per-token rates used to reconcile the published GPT-6 Astra run cost]. OpenAI developer documentation. Verified 2026-09-04. developers.openai.com
  • Terminal-Bench (2026). Terminal-Bench 3.0 [announcement; 74 tasks across 7 domains, per-version pass rates]. Verified 2026-08-29. tbench.ai
  • Terminal-Bench (2026). Continuous Benchmarks [semantic versioning policy for patch, minor and major releases]. Verified 2026-08-29. tbench.ai
  • Harbor Framework (2026). terminal-bench v4.0.0 [release notes naming the 8 removed tasks; task tree contains 66 directories]. GitHub. Verified 2026-08-29. github.com
  • Segato, G. (2026). Quantifying infrastructure noise in agentic coding evals [engineering research; 6 percentage-point gap across six resource configurations on Terminal-Bench 2.0]. Anthropic, February 5, 2026. Verified 2026-08-29. anthropic.com
  • Merrill, M. A., et al. (2026). Terminal-Bench: Benchmarking agents on hard, realistic tasks in command line interfaces [preprint]. arXiv:2601.11868. arxiv.org

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← Back to Models & benchmarks