Capital & Compute

Terminal-Bench

Benchmark· Coding & software-engineering agents· Checked 2026-09-04

Also known as Terminal-Bench 4.0, Terminal-Bench 3.0, Frontier-Bench, T-Bench

Terminal-Bench measures whether an AI agent can complete hard, realistic command-line work end to end inside a real terminal: building, configuring, training, debugging and securing systems. It is graded by running verification scripts against the actual end state, so an agent has to genuinely finish the job rather than produce a convincing transcript. The current release, version 4.0, contains 66 tasks and is led by Claude Opus 5 at 51.82%.

Key facts about the Terminal-Bench benchmark
What it measuresWhether an AI agent can complete hard, realistic command-line tasks (build, configure, train, debug, secure) end to end inside a real terminal.
Built byThe Terminal-Bench team (Marten, Shaw, Merrill) with the Laude Institute and 100+ community task contributors, 2026
Format66 human-verified containerized tasks (v4.0) across 7 domains, spanning software engineering, sysadmin, data science, ML, security and longer-horizon multi-container and GPU work
Scoring metricPass/fail per trial, graded by a verification script in a separate verifier container; each task is attempted 5 times, so a full board run is 330 trials
StatusActive
Representative top score58.2% (4.0) · GPT-6 Astra (Codex, max effort) · read 2026-09
Official leaderboardwww.tbench.ai/leaderboard/terminal-bench/4.0

How Terminal-Bench works

Version 4.0 contains 66 human-verified containerized tasks across 7 domains, from software engineering and sysadmin to data science, ML, security and longer-horizon work needing multi-container networks or GPUs. Each task is attempted 5 times, so a full leaderboard entry is 330 trials, and every task carries a flat 8-hour agent timeout. Grading is strictly pass or fail: the agent container and the verifier container are separate, and artifacts are downloaded from the agent and uploaded to the verifier at the end of a trial, which blocks many reward-hacking routes and lets results be re-graded when a verifier is fixed. Because the agent operates a real shell, a score reflects the model and its scaffolding together, not the model alone.

History and current status

Built by the Terminal-Bench team with the Laude Institute, it became the headline agentic benchmark of 2026 as SWE-bench Verified saturated. Its version history is the important part. Version 2.1 (May 2026) ran 89 tasks and topped out near 83%. Version 3.0 (2026-07-30) replaced the set with 74 harder tasks across 7 domains, developed under the working name Frontier-Bench, and the top score fell to 34.4%. Version 4.0 (2026-08-28) removed 8 tasks and fixed 19, leaving 66. Claude Opus 5 led it at 51.82% until GPT-6 Astra entered on its 2026-09-03 launch day and took the top score at 58.18% in Codex at max effort, with Claude Fable 5.1 and two lower Astra effort settings tied at 57.88% just behind; the board grew from 10 entries to 18 in the same week. From 3.0 onward it is governed by a published semantic-versioning policy: major releases change the environment or task set and force full re-runs, minor releases change verifiers and re-grade saved artifacts, patches change nothing that moves a score. Milestones for 4.1 (tamper-resistant verifiers) and 5.0 (new tasks) are open and unshipped as of September 2026.

What the score does not tell you

Scaffolding dependence is the main caveat: the same model scores differently under different harnesses, so a Terminal-Bench number is a statement about a stack rather than a model. Version churn is the second. A figure quoted without its version is close to meaningless, because 2.1, 3.0 and 4.0 differ by more than 30 points on the same name. The third is resolution: the published 95% confidence half-widths run from 2.62 to 3.85 points on the 4.0 board, and Anthropic research on infrastructure noise found a 6-point swing from resource configuration alone, recommending skepticism of any gap under 3 points. Pass-or-fail grading is honest but discards partial progress, which is noisier on a 66-task set than a larger one.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

Scored on those four axes, Terminal-Bench carries a concern score of 0 out of 8 and ranks 2 of 15, which puts it in the group worth quoting as it stands. See the reasoning behind that rating and how it compares with the rest of the field.

What Terminal-Bench scores actually mean

Terminal-Bench grades the end state of a machine, so a score is the share of jobs actually finished, not the share of plausible attempts. On v4.0 the leader reads 51.8%, meaning the best available agent fails roughly half of 66 verified tasks even with five attempts each. The version number is load-bearing and comparisons across it are invalid: the same board went from about 83% on v2.1 to 34.4% on v3.0 to 51.8% on v4.0, because each release swapped in a different task set under the same name. Treat any Terminal-Bench figure quoted without a version as unusable. The board also publishes total run cost, which spans $6.08 to $234.24 across the ten v4.0 entries, so cost per solved task differs by more than an order of magnitude between agents scoring similarly.

Who reports Terminal-Bench, and how to read it

Frontier labs quote it at launch, and tbench.ai runs the official board with grant support it discloses from OpenAI, Anthropic, Z.ai, SpaceX AI and the Laude Institute; Artificial Analysis has published independent runs of earlier versions. Always check the version, since 2.1, 3.0 and 4.0 are separate task sets, and check the harness, since the agent is named on every board row. Vendor-reported figures frequently lag the official board by a version. One practical trap: the version segment in a tbench.ai leaderboard URL is ignored, so requesting an older version returns the current rows.

When to weight Terminal-Bench in a model choice

Weight Terminal-Bench when the question is whether an agent can be trusted to complete work unattended, because pass/fail verification against the machine state is the closest thing in this directory to a production check. It is the right number for choosing a coding agent or harness, and a poor number for ranking a raw model, since the harness contributes heavily to the result. Always read the score together with the published run cost: the cost spread across entries is wider than the score spread, so the cheapest agent within a few points of the leader is usually the correct choice. Confirm the version before quoting, and confirm it again a month later, because the set updates continuously.

2026
First released
The Terminal-Bench team (Marten, Shaw, Merrill) with the Laude Institute and 100+ community task contributors
Active
Status today
As of September 12, 2026
58.2% (4.0)
Representative top score
Read 2026-09

Benchmarks to read alongside this one

Terminal-Bench: frequently asked questions

What is Terminal-Bench?
Terminal-Bench is a benchmark of hard, human-verified containerized command-line tasks covering software engineering, system administration, data science, machine learning and security. An agent works in a real terminal and is graded pass or fail by verification scripts that check the end state. Version 4.0 contains 66 tasks, each attempted five times.
Which model leads Terminal-Bench 4.0?
Claude Opus 5 on Claude Code at max effort leads with 51.82%, solving 171 of 330 trials, followed by Claude Fable 5 at 44.55% and GLM-5.3 at 41.82%. Confidence half-widths of 2.6 to 3.9 points mean closely spaced entries are not reliably separated.
Why do Terminal-Bench scores differ so much between versions?
Because the task set is replaced, not just tuned. Version 2.1 had 89 tasks and topped near 83%, version 3.0 introduced 74 harder tasks and the top score fell to 34.4%, and version 4.0 cut that to 66 tasks with a top score of 51.82%. A score is only comparable to another score on the same version.
Is Terminal-Bench better than SWE-bench?
It measures something different and is currently harder to contaminate. SWE-bench asks for a patch that passes tests; Terminal-Bench asks whether an agent can operate a real machine to reach a required end state, graded in a separate verifier container. For agentic work read Terminal-Bench; for issue resolution read SWE-bench Pro.

Sources

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →