SWE-bench Pro
SWE-bench Pro is the successor benchmark built for the period after SWE-bench Verified saturated. It uses longer, enterprise-grade software tasks, standardised agent scaffolding so results are comparable across models, and held-out splits that are not published, which makes contamination much harder.
| What it measures | Whether agents can solve long-horizon, enterprise-grade software-engineering tasks under standardized scaffolding, designed to resist contamination. |
|---|---|
| Built by | Scale AI (Scale Labs), 2025 |
| Format | 1,865 problems across 41 maintained repos, with public, held-out, and commercial (proprietary) splits; Python, Go, TypeScript and JavaScript |
| Scoring metric | % resolved (pass@1) under standardized agent scaffolding |
| Status | Active |
| Representative top score | 59.1% (public set) · GPT-5.4 (xHigh) · read 2026-06 |
| Official leaderboard | labs.scale.com/leaderboard/swe_bench_pro_public |
How SWE-bench Pro works
The set contains 1,865 problems across 41 maintained repositories in Python, Go, TypeScript and JavaScript, divided into a public split, a held-out split and a commercial split drawn from proprietary code. Every model runs under the same scaffolding rather than each lab’s own harness, which is the design decision that makes the leaderboard genuinely comparable. Scoring is the familiar percentage resolved on one attempt.
History and current status
Scale AI introduced it in 2025 as the answer to two problems at once: Verified had saturated, and cross-lab comparison had become meaningless because every lab reported under a different harness. Adoption accelerated in 2026 when OpenAI recommended reporting Pro instead of Verified. The public standardised leaderboard has GPT-5.4 at xHigh effort leading at 59.1% as of June 2026.
What the score does not tell you
The commercial split is proprietary, so nobody outside the operator can reproduce those results, and the benchmark is run by a company that also sells evaluation services. The gap between standardised board scores and self-reported figures is large: several 2026 launches claimed the mid-60s under their own scaffolding while the standardised board showed the high 50s. Both can be accurate; they are simply not the same measurement.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Who reports SWE-bench Pro, and how to read it
OpenAI now points to Pro rather than Verified, and most 2026 frontier launches quote a Pro number. Check whether the figure comes from the standardised public board or from the lab’s own scaffolding, because the two differ by several points and only the first is comparable across labs.
Benchmarks to read alongside this one
SWE-bench Verified
The same real-GitHub-issue resolution task as SWE-bench, restricted to a human-validated subset where the issue is solvable and the tests are not broken.
SWE-bench
Whether a system can resolve a real GitHub issue by generating a patch that passes the repository's hidden tests.
Terminal-Bench
Whether an AI agent can complete hard, realistic command-line tasks (build, configure, train, debug, secure) end to end inside a real terminal.
SWE-bench Pro: frequently asked questions
- What is SWE-bench Pro?
- SWE-bench Pro is a 2025 benchmark of 1,865 long-horizon, enterprise-grade software-engineering problems across 41 repositories in Python, Go, TypeScript and JavaScript. It runs every model under standardised scaffolding and keeps held-out and commercial splits unpublished to resist contamination.
- Why are SWE-bench Pro scores so much lower than SWE-bench Verified?
- Three reasons: the tasks are longer and harder, the held-out splits have not leaked into training data, and every model runs under the same scaffolding instead of a lab-tuned harness. Models near 80% on Verified typically land in the 45 to 60% range on Pro.
- Is SWE-bench Pro run independently?
- It is run by Scale AI rather than by a model vendor, which makes it more independent than a lab-run evaluation, but Scale also sells evaluation services and the commercial split is proprietary and not externally reproducible. Treat it as the best comparable public coding board rather than a neutral ground truth.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.