Capital & Compute
Benchmark· Coding & software-engineering agents· Checked 2026-07-10

Terminal-Bench

Also known as Terminal-Bench 2.0, T-Bench

Terminal-Bench measures whether an AI agent can complete hard, realistic command-line work end to end inside a real terminal: building, configuring, training, debugging and securing systems. It is graded by running verification scripts against the actual end state, so an agent has to genuinely finish the job rather than produce a convincing transcript.

Key facts about the Terminal-Bench benchmark
What it measuresWhether an AI agent can complete hard, realistic command-line tasks (build, configure, train, debug, secure) end to end inside a real terminal.
Built byStanford and the Laude Institute, 2026
Format89 human-verified containerized tasks (v2.0) spanning software engineering, sysadmin, data science, ML and security
Scoring metricPass/fail, graded by verification scripts in the agent's Docker environment (pass@1)
StatusActive
Representative top score83.4% (v2.1) · Codex (GPT-5.5) · read 2026-06
Official leaderboardwww.tbench.ai

How Terminal-Bench works

Version 2.0 contains 89 human-verified containerised tasks spanning software engineering, system administration, data science, machine learning and security. Each task runs in the agent’s own Docker environment, and grading is strictly pass or fail: a verification script inspects the resulting state. Because the agent operates a real shell, the reported score reflects the model and its scaffolding together, not the model alone.

History and current status

Built by Stanford and the Laude Institute, it became a headline agentic benchmark in 2026 as SWE-bench Verified saturated. The official board for v2.1 shows Codex CLI with GPT-5.5 at 83.4%, Claude Fable 5 at 83.1% and Claude Sonnet 5 self-reporting 80.4% as of June 2026. Artificial Analysis independent v2.1 run ranks GPT-5.6 Sol at 89.5%, ahead of the official board figures. The same team went on to release Harbor-Index.

What the score does not tell you

The scaffolding dependence is the main caveat: the same model scores differently under different harnesses, so a Terminal-Bench number is a statement about a stack rather than a model. Pass-or-fail grading is a strength for honesty but discards partial progress, which makes the metric noisy on a set of only 89 tasks. Independent and official runs also disagree by several points.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

Who reports Terminal-Bench, and how to read it

Frontier labs quote it at launch, and both tbench.ai and Artificial Analysis publish independent results. Check the version, since v2.0 and v2.1 differ, and check whether the figure is from the official board or a third-party run, because in 2026 those have not agreed.

2026
First released
Stanford and the Laude Institute
Active
Status today
As of July 27, 2026
83.4% (v2.1)
Representative top score
Read 2026-06

Benchmarks to read alongside this one

Terminal-Bench: frequently asked questions

What is Terminal-Bench?
Terminal-Bench is a benchmark of 89 human-verified containerised command-line tasks covering software engineering, system administration, data science, machine learning and security. An agent works in a real terminal and is graded pass or fail by verification scripts that check the end state.
Why do Terminal-Bench scores differ between sources?
Because the score measures a model plus its agent scaffolding. Different harnesses produce different results for the same model, and the official tbench.ai board and Artificial Analysis independent runs have reported figures several points apart for the same version.
Is Terminal-Bench better than SWE-bench?
It measures something different and is currently harder to contaminate. SWE-bench asks for a patch that passes tests; Terminal-Bench asks whether an agent can operate a real machine to reach a required end state. For agentic work, read Terminal-Bench; for issue resolution, read SWE-bench Pro.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory