Capital & Compute

SWE-bench Verified

Benchmark· Coding & software-engineering agents· Checked 2026-06-29

SWE-bench Verified is a 500-task subset of SWE-bench in which humans checked every task to confirm that the issue is actually solvable and the tests are not broken. It became the most-quoted coding benchmark in the industry because it removed the noise in the original set, and it is now the most-quoted saturated one.

Key facts about the SWE-bench Verified benchmark
What it measuresThe same real-GitHub-issue resolution task as SWE-bench, restricted to a human-validated subset where the issue is solvable and the tests are not broken.
Built byOpenAI (with the SWE-bench authors), 2024
Format500 human-validated task instances drawn from SWE-bench Full (Python)
Scoring metric% resolved (pass@1)
StatusSaturated
Representative top score~95% · Claude Fable 5 · read 2026-06
Official leaderboardwww.swebench.com

How SWE-bench Verified works

The task is identical to SWE-bench: read a real GitHub issue, patch the repository, pass the hidden tests. The difference is curation. Professional developers reviewed candidate tasks and kept only those where the issue statement was sufficient to produce the fix and the test suite correctly distinguished a fix from a non-fix. The result is 500 tasks where a failure is much more likely to be the model’s fault than the benchmark’s fault.

History and current status

OpenAI released Verified in 2024, working with the original SWE-bench authors, after analysis showed a substantial fraction of the full set was unsolvable as specified. It rapidly displaced the full set in model cards. Scores climbed from the low tens of percent to near the ceiling within roughly two years. In February 2026 OpenAI said it had stopped reporting Verified, citing broken tests and training-data exposure, and recommended the harder SWE-bench Pro.

What the score does not tell you

Verified fixed the task-quality problem but not the contamination problem: the underlying repositories and their fix commits remain public. When frontier models bunch near the top, the remaining spread is scaffolding and luck rather than capability. That is why a very high Verified score is now better read as a contamination and saturation signal than as evidence of engineering skill.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, why benchmark saturation makes a high score meaningless.

Scored on those four axes, SWE-bench Verified carries a concern score of 7 out of 8 and ranks 13 of 15, which puts it in the group to stop quoting. See the reasoning behind that rating and how it compares with the rest of the field.

What SWE-bench Verified scores actually mean

The usable range on Verified closed years ago. In 2024 a frontier model landed in the 30s to 40s; by 2025 the leaders were in the 60s and 70s; by 2026 they cluster above 80% with the top reading near 95%. Inside that top band the spread is not capability. Re-running the same model under a different agent harness moves the number by ten points or more, so a two-point gap between two labs' reported figures carries no information about which model is better at fixing code. Below roughly 40% the score still discriminates, which is why Verified remains a reasonable smoke test for a small or open-weight model, and useless for ranking the frontier.

Who reports SWE-bench Verified, and how to read it

Verified was the default coding row in launch tables from 2024 through 2026, usually reported under each lab’s own agent harness, which makes cross-lab comparison unreliable. Its own originator has now stepped away from it. Where a lab still quotes it, check whether a Pro or Terminal-Bench number is reported alongside, and weight those instead.

When to weight SWE-bench Verified in a model choice

Use Verified to answer one question: can this model do repository-scale editing at all? That is a real question for a 7B local model or a cheap API tier, and Verified answers it cheaply. Do not use it to choose between frontier models, and treat a launch table that leads with Verified in 2026 as a soft signal that the harder numbers were less flattering. The organisation that built the set stopped reporting it in February 2026 and pointed at SWE-bench Pro instead, which is the strongest possible statement about its remaining value. Where a lab quotes Verified, look for a Pro or Terminal-Bench figure in the same table and weight that one.

2024
First released
OpenAI (with the SWE-bench authors)
Saturated
Status today
As of September 12, 2026
~95%
Representative top score
Read 2026-06

Benchmarks to read alongside this one

SWE-bench Verified: frequently asked questions

What is the difference between SWE-bench and SWE-bench Verified?
SWE-bench is the original 2,294-task set. SWE-bench Verified is a 500-task subset that OpenAI and the SWE-bench authors hand-checked in 2024 to remove tasks with broken tests or issue descriptions too vague to solve. Verified is cleaner and became the version labs actually report.
Why did OpenAI stop reporting SWE-bench Verified?
OpenAI said in February 2026 that it had found broken tests and evidence of training-data exposure in the set, and now recommends SWE-bench Pro instead. In other words the organisation that created Verified concluded its scores had stopped tracking real capability.
Is SWE-bench Verified saturated?
Yes. Frontier models cluster near the ceiling, so differences between the leaders fall inside the noise created by different agent scaffolding. When a benchmark can no longer rank the strongest systems, it is measuring its own ceiling rather than the models.

Sources

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →