Which AI Benchmarks Are Worth Trusting
Every model launch quotes benchmark scores, and most of those scores no longer mean what the number implies. This page ranks the 15 most-quoted AI benchmarks by how much reason there is to distrust them, on four axes that can be checked against public evidence. For what each benchmark is and what it currently scores, see thefull AI benchmarks directory. For how scores get inflated in the first place, seehow benchmark scores get gamed.
Which AI benchmarks should you trust in 2026?
6 of the 15 most-quoted AI benchmarks clear the bar, ranked most trustworthy first: SWE-bench Pro, Terminal-Bench 4, LiveBench, OSWorld 2.0, tau-bench and METR Time Horizon. Each keeps its test data held back or rotating, sits well clear of its ceiling, and grades by running code rather than comparing text. SWE-bench Verified, HumanEval and MMLU fail on all three counts.
The three groups
A benchmark earns its place by what it holds back, not by how hard it looks. Each group links to its rows below.
- 6 of 15Worth trustingHeld-out or rotating test data, scores well clear of the ceiling, and grading that runs code rather than comparing text. Quote these numbers.SWE-bench Pro, Terminal-Bench 4, LiveBench, OSWorld 2.0, tau-bench, METR Time Horizon.
- 6 of 15Use with the caveatEach measures something real, and each carries one problem big enough that the number needs a sentence of context beside it. Never quote these bare.ARC-AGI-3, FrontierMath, ARC-AGI-2, GPQA Diamond, LMArena, MMLU-Pro.
- 3 of 15Stop quotingPublic test data, scores bunched at the ceiling, and a documented reason the field has already moved on. A high score here is evidence of contamination, not capability.SWE-bench Verified, HumanEval, MMLU.
| Rank | Benchmark | Contamination | Saturation | Gameability | Real-world gap | Concern | Verdict |
|---|---|---|---|---|---|---|---|
| 1 | SWE-bench Pro (real GitHub fixes) | Low | Low | Low | Low | 0 of 8 | trust |
| 2 | Terminal-Bench 4 (terminal tasks) | Low | Low | Low | Low | 0 of 8 | trust |
| 3 | LiveBench (contamination-limited composite) | Low | Low | Low | Medium | 1 of 8 | trust |
| 4 | OSWorld 2.0 (computer use) | Low | Low | Medium | Low | 1 of 8 | trust |
| 5 | tau-bench (tool use under policy) | Medium | Low | Low | Low | 1 of 8 | trust |
| 6 | METR Time Horizon (task length, not accuracy) | Low | Low | Medium | Medium | 2 of 8 | trust |
| 7 | ARC-AGI-3 (interactive reasoning) | Low | Low | Medium | High | 3 of 8 | situational |
| 8 | FrontierMath (research math) | Medium | High | Low | Medium | 4 of 8 | situational |
| 9 | ARC-AGI-2 (abstract reasoning) | Low | High | Medium | High | 5 of 8 | situational |
| 10 | GPQA Diamond (PhD-level science) | Medium | High | Medium | Medium | 5 of 8 | situational |
| 11 | LMArena (human preference) | Medium | Low | High | High | 5 of 8 | situational |
| 12 | MMLU-Pro (knowledge MCQ, harder) | Medium | Medium | Medium | High | 5 of 8 | situational |
| 13 | SWE-bench Verified (real GitHub fixes) | High | High | High | Medium | 7 of 8 | stop-quoting |
| 14 | HumanEval (function-level coding) | High | High | High | High | 8 of 8 | stop-quoting |
| 15 | MMLU (knowledge MCQ) | High | High | High | High | 8 of 8 | stop-quoting |
| # | Benchmark | Contamination | Saturation | Gameability | Real-world gap | Concern | Verdict |
|---|---|---|---|---|---|---|---|
| 1 | SWE-bench Proreal GitHub fixes | Low | Low | Low | Low | 0 / 8 | Trust |
| 2 | Terminal-Bench 4terminal tasks | Low | Low | Low | Low | 0 / 8 | Trust |
| 3 | LiveBenchcontamination-limited composite | Low | Low | Low | Medium | 1 / 8 | Trust |
| 4 | OSWorld 2.0computer use | Low | Low | Medium | Low | 1 / 8 | Trust |
| 5 | tau-benchtool use under policy | Medium | Low | Low | Low | 1 / 8 | Trust |
| 6 | METR Time Horizontask length, not accuracy | Low | Low | Medium | Medium | 2 / 8 | Trust |
| 7 | ARC-AGI-3interactive reasoning | Low | Low | Medium | High | 3 / 8 | Caveat |
| 8 | FrontierMathresearch math | Medium | High | Low | Medium | 4 / 8 | Caveat |
| 9 | ARC-AGI-2abstract reasoning | Low | High | Medium | High | 5 / 8 | Caveat |
| 10 | GPQA DiamondPhD-level science | Medium | High | Medium | Medium | 5 / 8 | Caveat |
| 11 | LMArenahuman preference | Medium | Low | High | High | 5 / 8 | Caveat |
| 12 | MMLU-Proknowledge MCQ, harder | Medium | Medium | Medium | High | 5 / 8 | Caveat |
| 13 | SWE-bench Verifiedreal GitHub fixes | High | High | High | Medium | 7 / 8 | Stop |
| 14 | HumanEvalfunction-level coding | High | High | High | High | 8 / 8 | Stop |
| 15 | MMLUknowledge MCQ | High | High | High | High | 8 / 8 | Stop |
The four axes, and why they sum
A benchmark score is a claim about a model. Four things can break that claim, and they break it in different ways, so each is rated separately and then added. All four are scored so that higher always means worse, which is the only reason a single concern number out of 8 is meaningful rather than an average of incompatible things.
- Contamination. How likely the test data leaked into training. Higher = worse.
- Saturation. How close top models sit to the ceiling, so the score stops telling models apart. Higher = worse.
- Gameability. How easily a score is inflated without real capability. Higher = worse.
- Real-world gap. Gap between the headline score and real task performance. Higher = worse.
The bands are an editorial line, stated so you can disagree with it precisely. A benchmark carrying at most one Medium and nothing High is in the trust group. Anything from three to five is situational: it measures something real and carries one problem large enough that the number needs a sentence beside it. Six or more, which in practice means two or more High ratings, is in the stop-quoting group.
What the axes deliberately do not measure is difficulty. A hard benchmark is not a trustworthy one, and several of the hardest here sit mid-table: ARC-AGI-2 is brutal and nearly contamination-proof, and it still lands in the situational group because the headroom is gone and the thing it measures is fluid reasoning rather than work anyone is paying for.
The 6 benchmarks worth trusting
The pattern in this group is not cleverness. It is what each one refuses to publish, and what it does instead of asking a model to pick from a list. Most of them grade by executing something rather than by comparing text, the rest keep their questions moving, and none of them is close to a ceiling. The top two tie with nothing flagged at all.
1. SWE-bench Pro
real GitHub fixes. Concern 0 of 8. The replacement OpenAI itself now recommends over Verified. Held-out and commercial splits keep the answers off GitHub, and standardized scaffolding removes the harness inflation that lets vendors publish their own numbers. The effect is blunt: models near 80% on Verified land in the 45 to 60% range here, which is the size of the contamination and scaffold premium the older benchmark was hiding. Full SWE-bench Pro profile.
2. Terminal-Bench 4
terminal tasks. Concern 0 of 8. Graded pass/fail by a verifier running in a separate container from the agent, against real test suites in a real terminal, which is hard to fake. The task set is written for the benchmark by a community of 100+ contributors and reviewers rather than scraped from public issues, so there is nothing public to memorize. Saturation is actively managed rather than tolerated: v4.0 deleted tasks every current-generation model solved 5/5. Two real caveats. Versions are not comparable, since v2.1, v3.0 and v4.0 are different task sets sharing a name and differ by more than 30 points. And leakage rises as each public set ages, which is why the operators warn against training on it. Full Terminal-Bench 4 profile.
3. LiveBench
contamination-limited composite. Concern 1 of 8. The strongest composite alternative to preference arenas: no LLM judge, no human votes, and monthly question rotation that limits what can be memorized. Its authors renamed it from contamination-free to contamination-limited, which is the honest framing and the reason it scores well here. Real-world gap is medium because the task mix is academic rather than production work. Full LiveBench profile.
4. OSWorld 2.0
computer use. Concern 1 of 8. Exists because agents reached 83.5% on OSWorld 1.0. Real desktop tasks in a real OS, where a skilled human needs a median of roughly 1.6 hours per task, and the headroom is enormous. Marked medium on gaming for one reason: it reports a partial score alongside full completion, and vendors quote the partial figure from their own harnesses. Full OSWorld 2.0 profile.
5. tau-bench
tool use under policy. Concern 1 of 8. Built to expose unreliability rather than to be topped. Scoring uses pass^k, which requires the same task to succeed across k independent trials, so a lucky single run earns nothing: pass^8 falls far below pass^1 for every model tested. The public task set has been available since 2024, which is the one real contamination exposure here. Full tau-bench profile.
6. METR Time Horizon
task length, not accuracy. Concern 2 of 8. The only entry whose unit is time rather than a percentage, so it cannot saturate the way a capped score does: it reports the task length a model completes at 50% success. The launch paper found a doubling roughly every seven months since 2019. Marked medium on gaming and real-world gap because the fitted horizon is a derived statistic that moves with the task mix and the harness, not a directly observed result. Full METR Time Horizon profile.
The 6 that need a caveat
None of these is junk, and quoting any of them bare is still misleading. Each one has a single dominant problem, and knowing which problem it is tells you what the number can and cannot support.
7. ARC-AGI-3
interactive reasoning. Concern 3 of 8. Where the ARC gap moved. Interactive game environments with no instructions, launched March 2026 with humans at 100% and frontier AI at 0.51%, and still far from solved. Scoring on human-level action efficiency fences brute-force search better than v2 did. The real-world gap is intentional: it measures fluid reasoning, not production work. Harness choice is load-bearing, which is the gaming exposure. Full ARC-AGI-3 profile.
8. FrontierMath
research math. Concern 4 of 8. Held back from the public and from most labs, and designed against guessing, so it resists gaming better than anything else here. The contamination rating is not about leakage to the field: Epoch discloses that OpenAI commissioned and funds it, owns the problems, and has problem and solution access outside a held-out set. That is a conflict of interest for one lab specifically. Now saturated on the graded tiers, leaving the Open Problems track as the frontier. Full FrontierMath profile.
9. ARC-AGI-2
abstract reasoning. Concern 5 of 8. Private and semi-private eval sets keep contamination low, but the headroom is gone. Verified scores went from 54% in December 2025 to past the 85% target within seven months. Gameable by high-compute brute force, fenced by cost-per-task caps rather than by design. The real-world gap is intentional: it measures fluid reasoning, not production work. Full ARC-AGI-2 profile.
10. GPQA Diamond
PhD-level science. Concern 5 of 8. Google-proof PhD science questions with a 65% expert baseline and 34% for skilled non-experts with web access. The reasoning-gated design resists pure retrieval far better than MMLU, but top models are now past the human-expert baseline, so it has largely saturated. Only 198 items, which means a handful of questions swings the headline number. Full GPQA Diamond profile.
11. LMArena
human preference. Concern 5 of 8. Measures preference, not capability, and style is load-bearing: length, formatting and tone move votes. Meta topped it in April 2025 with a chat-tuned variant that was not the released model, and the operator said afterwards that this violated its expectations. The Leaderboard Illusion paper documents how privileged private testing and selective disclosure let large labs overfit the ranking. Full LMArena profile.
12. MMLU-Pro
knowledge MCQ, harder. Concern 5 of 8. Built to replace saturated MMLU: ten answer options instead of four and reasoning-heavy items drop scores 16 to 33 points and separate frontier models again. It is the right substitute if a knowledge score is required, but it is still public multiple choice, and by mid-2026 the top tier is compressing again, so it is walking the same path. Full MMLU-Pro profile.
The 3 benchmarks to stop quoting
These three share a profile: public test data, top scores bunched at the ceiling, and a documented reason the field has already moved on. A record on any of them is better read as evidence of contamination than of capability, and the most striking case is the one that is still quoted most.
13. SWE-bench Verified
real GitHub fixes. Concern 7 of 8. The most-quoted coding number in launch posts, and OpenAI stopped reporting it in February 2026 because frontier models reproduce the reference patches. The gold patches are public on GitHub. SWE-Bench+ found answer leakage and weak tests. Top models sit near the ceiling. Superseded in practice by SWE-bench Pro and by execution-graded successors, and still quoted anyway. Full SWE-bench Verified profile.
14. HumanEval
function-level coding. Concern 8 of 8. Effectively solved: frontier models sit near 99% pass@1, so it cannot separate strong models at all. 164 short standalone functions with public tests, present in essentially every code training corpus. It survives as a historical baseline and as a number to pad a launch table with, not as evidence about engineering ability. Full HumanEval profile.
15. MMLU
knowledge MCQ. Concern 8 of 8. Fully public four-option multiple choice, and demonstrably memorized: GPT-4 reconstructed masked answer options 57% of the time in a peer-reviewed contamination study. Top models bunch in the low-to-mid 90s, so rank differences are mostly noise. The field already built its replacements, MMLU-Pro and MMLU-Redux, and still quotes the original. Full MMLU profile.
What to quote instead
The substitutions are close to one-for-one, which removes the usual excuse for leaning on a saturated number. For coding, SWE-bench Pro replaces SWE-bench Verified and the same models drop from around 80 percent into the 45 to 60 percent range, which is the size of the premium the old benchmark was hiding. For knowledge, MMLU-Pro replaces MMLU and takes 16 to 33 points off with ten answer options instead of four, though it is walking the same road and will need replacing too. For function-level coding, nothing replaces HumanEval, because nothing needs to: it measures whether a model can write a short standalone function, and every frontier model can.
For agentic work the honest answer is that there is no single number. Terminal-Bench 4 covers terminal tasks, tau-bench covers tool use under a policy, and OSWorld 2.0 covers driving a desktop. They measure different jobs and disagree about which model is best, which is information rather than a defect. Thebenchmarks directory carries the current record on each, and the value leaderboard puts capability against what it costs to obtain.
One habit does more than any substitution: quote the version. Terminal-Bench v2.1, v3.0 and v4.0 are different task sets sharing a name, and the top score differs across them by more than 30 points. A benchmark number without a version is not a claim anyone can check, and versionless comparisons are the most common way an honest score becomes a misleading one.
How this ranking was made, and what it cannot tell you
The ratings are judgement, not measurement, and the distinction matters enough to state plainly. No score in the table above was produced by this site, nothing here is hands-on tested, and each of the four axes is scored from evidence that is public and linkable: whether the test set is published, whether frontier scores have bunched, whether the funding arrangement gives one lab an advantage, and whether the harness moves the result. Every row links the primary source its rating rests on, so a disagreement can be argued at the evidence rather than about the conclusion.
Three limits are worth naming. Concern scores are ordinal, so a benchmark at 4 is not twice as compromised as one at 2. The axes are unweighted, and there is a real argument that contamination should count for more than real-world gap on a page about whether to believe a number. And the ranking moves faster than it looks: ARC-AGI-2 went from 54 percent in December 2025 to past its 85 percent target within seven months, which would have moved it a whole band in that time. Current headline scores are read from thebenchmarks directory rather than copied onto this page, so a score that moves cannot leave a contradicting duplicate behind, but the trust ratings themselves need re-checking as the field saturates its own tests.
Frequently asked questions
- Which AI benchmark should you trust for coding?
- SWE-bench Pro and Terminal-Bench 4, which tie at the top with nothing flagged on any of the four axes. Both execute code and grade pass or fail against real test suites, both keep their answers off public GitHub, and both sit far from their ceiling. Choose between them by the job: Pro for repository-level issue fixing, Terminal-Bench for whole-task work in a shell. SWE-bench Verified is the number you will see quoted instead, and it is the one benchmark here whose most prominent user stopped reporting it.
- Why is SWE-bench Verified still quoted if OpenAI stopped reporting it?
- Because it is legible and the alternatives are not yet. Verified has years of published scores behind it, so a launch post can show a number that readers recognize and compare. SWE-bench Pro drops the same models from around 80 percent into the 45 to 60 percent range, which is a worse headline for an identical model. The incentive runs toward the older number, not the better one.
- What does benchmark contamination actually mean?
- That the test questions, or the answers, ended up in the training data, so the model is partly recalling rather than reasoning. It is measurable: a peer-reviewed study found GPT-4 could reconstruct masked MMLU answer options 57 percent of the time. Once a public test set is a few years old, contamination is the default assumption, not the exception.
- What does saturation mean for a launch announcement?
- That the benchmark has stopped separating models, so a new record on it is not evidence of much. When every frontier model scores in the low-to-mid 90s on MMLU or near 99 percent on HumanEval, the remaining spread is closer to noise than to capability. A saturated benchmark in a launch table is padding.
- Can a vendor game a benchmark score?
- Yes, and the documented cases are not subtle. Meta topped LMArena in April 2025 with a chat-tuned variant that was not the model it released, and the operator said afterwards that this violated its expectations. More routinely, vendors publish results from their own harness rather than the public board, which is why scaffold-sensitive benchmarks are marked as gameable here even when the task set is clean.
- Which AI benchmarks are private or held back?
- FrontierMath is held back from the public and from most labs, and the ARC-AGI sets keep semi-private and private splits. Holding data back is the strongest single defence against contamination, and it costs something: you cannot audit what you cannot see, and FrontierMath carries a disclosed funding conflict with one lab that has problem and solution access.
- Are these ratings measured or judged?
- Judged, from documented properties, and the page says so rather than implying a measurement. Each of the four axes is scored Low, Medium or High on evidence that is public and linkable: whether the test set is published, whether top scores have bunched at the ceiling, who funds the benchmark, and whether the harness moves the result. Nothing here is hands-on tested, and no score in the table was produced by this site.
- How current is this ranking?
- Every row was re-verified against its primary source in September 2026, and current headline scores are read from the benchmarks directory rather than copied here, so a score that moves cannot leave a stale duplicate behind. Benchmarks saturate fast: ARC-AGI-2 went from 54 percent to past its 85 percent target inside seven months. Check the linked source before quoting any figure.
Sources
Every row above links the primary source its rating rests on. The evidence the four axes lean on hardest:
- Deng, Zhao, Tang, Gerstein and Cohan (2024). Investigating Data Contamination in Modern Benchmarks for Large Language Models. NAACL 2024 (peer-reviewed).aclanthology.org
- OpenAI (2026). Why we no longer evaluate SWE-bench Verified. Vendor statement.openai.com
- Singh et al. (2025). The Leaderboard Illusion. Preprint.arxiv.org/abs/2504.20879
- Akhtar et al. (2026). When AI Benchmarks Plateau. Preprint.arxiv.org/abs/2602.16763
- Epoch AI (2025). Clarifying the creation and use of FrontierMath. Benchmark operator statement.epoch.ai
- Terminal-Bench (2026). Terminal-Bench 3.0. Benchmark operator announcement.tbench.ai
- Rein et al. (2023). GPQA: A Graduate-Level Google-Proof Q&A Benchmark. Preprint.arxiv.org/abs/2311.12022