Terminal-Bench
Also known as Terminal-Bench 2.0, T-Bench
Terminal-Bench measures whether an AI agent can complete hard, realistic command-line work end to end inside a real terminal: building, configuring, training, debugging and securing systems. It is graded by running verification scripts against the actual end state, so an agent has to genuinely finish the job rather than produce a convincing transcript.
| What it measures | Whether an AI agent can complete hard, realistic command-line tasks (build, configure, train, debug, secure) end to end inside a real terminal. |
|---|---|
| Built by | Stanford and the Laude Institute, 2026 |
| Format | 89 human-verified containerized tasks (v2.0) spanning software engineering, sysadmin, data science, ML and security |
| Scoring metric | Pass/fail, graded by verification scripts in the agent's Docker environment (pass@1) |
| Status | Active |
| Representative top score | 83.4% (v2.1) · Codex (GPT-5.5) · read 2026-06 |
| Official leaderboard | www.tbench.ai |
How Terminal-Bench works
Version 2.0 contains 89 human-verified containerised tasks spanning software engineering, system administration, data science, machine learning and security. Each task runs in the agent’s own Docker environment, and grading is strictly pass or fail: a verification script inspects the resulting state. Because the agent operates a real shell, the reported score reflects the model and its scaffolding together, not the model alone.
History and current status
Built by Stanford and the Laude Institute, it became a headline agentic benchmark in 2026 as SWE-bench Verified saturated. The official board for v2.1 shows Codex CLI with GPT-5.5 at 83.4%, Claude Fable 5 at 83.1% and Claude Sonnet 5 self-reporting 80.4% as of June 2026. Artificial Analysis independent v2.1 run ranks GPT-5.6 Sol at 89.5%, ahead of the official board figures. The same team went on to release Harbor-Index.
What the score does not tell you
The scaffolding dependence is the main caveat: the same model scores differently under different harnesses, so a Terminal-Bench number is a statement about a stack rather than a model. Pass-or-fail grading is a strength for honesty but discards partial progress, which makes the metric noisy on a set of only 89 tasks. Independent and official runs also disagree by several points.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Who reports Terminal-Bench, and how to read it
Frontier labs quote it at launch, and both tbench.ai and Artificial Analysis publish independent results. Check the version, since v2.0 and v2.1 differ, and check whether the figure is from the official board or a third-party run, because in 2026 those have not agreed.
Benchmarks to read alongside this one
SWE-bench Pro
Whether agents can solve long-horizon, enterprise-grade software-engineering tasks under standardized scaffolding, designed to resist contamination.
OSWorld 2.0
Whether a computer-use agent can finish long-horizon professional work on a real desktop, coordinating several applications and leaving the machine in the correct final state.
Aider Polyglot
How well a model writes and correctly edits code across many languages, including applying diffs in the right format and self-correcting after test failures.
Terminal-Bench: frequently asked questions
- What is Terminal-Bench?
- Terminal-Bench is a benchmark of 89 human-verified containerised command-line tasks covering software engineering, system administration, data science, machine learning and security. An agent works in a real terminal and is graded pass or fail by verification scripts that check the end state.
- Why do Terminal-Bench scores differ between sources?
- Because the score measures a model plus its agent scaffolding. Different harnesses produce different results for the same model, and the official tbench.ai board and Artificial Analysis independent runs have reported figures several points apart for the same version.
- Is Terminal-Bench better than SWE-bench?
- It measures something different and is currently harder to contaminate. SWE-bench asks for a patch that passes tests; Terminal-Bench asks whether an agent can operate a real machine to reach a required end state. For agentic work, read Terminal-Bench; for issue resolution, read SWE-bench Pro.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.