SWE-bench Verified
SWE-bench Verified is a 500-task subset of SWE-bench in which humans checked every task to confirm that the issue is actually solvable and the tests are not broken. It became the most-quoted coding benchmark in the industry because it removed the noise in the original set, and it is now the most-quoted saturated one.
| What it measures | The same real-GitHub-issue resolution task as SWE-bench, restricted to a human-validated subset where the issue is solvable and the tests are not broken. |
|---|---|
| Built by | OpenAI (with the SWE-bench authors), 2024 |
| Format | 500 human-validated task instances drawn from SWE-bench Full (Python) |
| Scoring metric | % resolved (pass@1) |
| Status | Saturated |
| Representative top score | ~95% · Claude Fable 5 · read 2026-06 |
| Official leaderboard | www.swebench.com |
How SWE-bench Verified works
The task is identical to SWE-bench: read a real GitHub issue, patch the repository, pass the hidden tests. The difference is curation. Professional developers reviewed candidate tasks and kept only those where the issue statement was sufficient to produce the fix and the test suite correctly distinguished a fix from a non-fix. The result is 500 tasks where a failure is much more likely to be the model’s fault than the benchmark’s fault.
History and current status
OpenAI released Verified in 2024, working with the original SWE-bench authors, after analysis showed a substantial fraction of the full set was unsolvable as specified. It rapidly displaced the full set in model cards. Scores climbed from the low tens of percent to near the ceiling within roughly two years. In February 2026 OpenAI said it had stopped reporting Verified, citing broken tests and training-data exposure, and recommended the harder SWE-bench Pro.
What the score does not tell you
Verified fixed the task-quality problem but not the contamination problem: the underlying repositories and their fix commits remain public. When frontier models bunch near the top, the remaining spread is scaffolding and luck rather than capability. That is why a very high Verified score is now better read as a contamination and saturation signal than as evidence of engineering skill.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, why benchmark saturation makes a high score meaningless.
Who reports SWE-bench Verified, and how to read it
Verified was the default coding row in launch tables from 2024 through 2026, usually reported under each lab’s own agent harness, which makes cross-lab comparison unreliable. Its own originator has now stepped away from it. Where a lab still quotes it, check whether a Pro or Terminal-Bench number is reported alongside, and weight those instead.
Benchmarks to read alongside this one
SWE-bench
Whether a system can resolve a real GitHub issue by generating a patch that passes the repository's hidden tests.
SWE-bench Pro
Whether agents can solve long-horizon, enterprise-grade software-engineering tasks under standardized scaffolding, designed to resist contamination.
Terminal-Bench
Whether an AI agent can complete hard, realistic command-line tasks (build, configure, train, debug, secure) end to end inside a real terminal.
SWE-bench Verified: frequently asked questions
- What is the difference between SWE-bench and SWE-bench Verified?
- SWE-bench is the original 2,294-task set. SWE-bench Verified is a 500-task subset that OpenAI and the SWE-bench authors hand-checked in 2024 to remove tasks with broken tests or issue descriptions too vague to solve. Verified is cleaner and became the version labs actually report.
- Why did OpenAI stop reporting SWE-bench Verified?
- OpenAI said in February 2026 that it had found broken tests and evidence of training-data exposure in the set, and now recommends SWE-bench Pro instead. In other words the organisation that created Verified concluded its scores had stopped tracking real capability.
- Is SWE-bench Verified saturated?
- Yes. Frontier models cluster near the ceiling, so differences between the leaders fall inside the noise created by different agent scaffolding. When a benchmark can no longer rank the strongest systems, it is measuring its own ceiling rather than the models.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.