Capital & Compute
Benchmark· Coding & software-engineering agents· Checked 2026-06-29

SWE-bench Verified

SWE-bench Verified is a 500-task subset of SWE-bench in which humans checked every task to confirm that the issue is actually solvable and the tests are not broken. It became the most-quoted coding benchmark in the industry because it removed the noise in the original set, and it is now the most-quoted saturated one.

Key facts about the SWE-bench Verified benchmark
What it measuresThe same real-GitHub-issue resolution task as SWE-bench, restricted to a human-validated subset where the issue is solvable and the tests are not broken.
Built byOpenAI (with the SWE-bench authors), 2024
Format500 human-validated task instances drawn from SWE-bench Full (Python)
Scoring metric% resolved (pass@1)
StatusSaturated
Representative top score~95% · Claude Fable 5 · read 2026-06
Official leaderboardwww.swebench.com

How SWE-bench Verified works

The task is identical to SWE-bench: read a real GitHub issue, patch the repository, pass the hidden tests. The difference is curation. Professional developers reviewed candidate tasks and kept only those where the issue statement was sufficient to produce the fix and the test suite correctly distinguished a fix from a non-fix. The result is 500 tasks where a failure is much more likely to be the model’s fault than the benchmark’s fault.

History and current status

OpenAI released Verified in 2024, working with the original SWE-bench authors, after analysis showed a substantial fraction of the full set was unsolvable as specified. It rapidly displaced the full set in model cards. Scores climbed from the low tens of percent to near the ceiling within roughly two years. In February 2026 OpenAI said it had stopped reporting Verified, citing broken tests and training-data exposure, and recommended the harder SWE-bench Pro.

What the score does not tell you

Verified fixed the task-quality problem but not the contamination problem: the underlying repositories and their fix commits remain public. When frontier models bunch near the top, the remaining spread is scaffolding and luck rather than capability. That is why a very high Verified score is now better read as a contamination and saturation signal than as evidence of engineering skill.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, why benchmark saturation makes a high score meaningless.

Who reports SWE-bench Verified, and how to read it

Verified was the default coding row in launch tables from 2024 through 2026, usually reported under each lab’s own agent harness, which makes cross-lab comparison unreliable. Its own originator has now stepped away from it. Where a lab still quotes it, check whether a Pro or Terminal-Bench number is reported alongside, and weight those instead.

2024
First released
OpenAI (with the SWE-bench authors)
Saturated
Status today
As of July 27, 2026
~95%
Representative top score
Read 2026-06

Benchmarks to read alongside this one

SWE-bench Verified: frequently asked questions

What is the difference between SWE-bench and SWE-bench Verified?
SWE-bench is the original 2,294-task set. SWE-bench Verified is a 500-task subset that OpenAI and the SWE-bench authors hand-checked in 2024 to remove tasks with broken tests or issue descriptions too vague to solve. Verified is cleaner and became the version labs actually report.
Why did OpenAI stop reporting SWE-bench Verified?
OpenAI said in February 2026 that it had found broken tests and evidence of training-data exposure in the set, and now recommends SWE-bench Pro instead. In other words the organisation that created Verified concluded its scores had stopped tracking real capability.
Is SWE-bench Verified saturated?
Yes. Frontier models cluster near the ceiling, so differences between the leaders fall inside the noise created by different agent scaffolding. When a benchmark can no longer rank the strongest systems, it is measuring its own ceiling rather than the models.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory