Capital & Compute

Which AI Benchmarks Are Worth Trusting

Ranking· Updated September 2026· 15 benchmarks· 6 pass

Every model launch quotes benchmark scores, and most of those scores no longer mean what the number implies. This page ranks the 15 most-quoted AI benchmarks by how much reason there is to distrust them, on four axes that can be checked against public evidence. For what each benchmark is and what it currently scores, see thefull AI benchmarks directory. For how scores get inflated in the first place, seehow benchmark scores get gamed.

Which AI benchmarks should you trust in 2026?

6 of the 15 most-quoted AI benchmarks clear the bar, ranked most trustworthy first: SWE-bench Pro, Terminal-Bench 4, LiveBench, OSWorld 2.0, tau-bench and METR Time Horizon. Each keeps its test data held back or rotating, sits well clear of its ceiling, and grades by running code rather than comparing text. SWE-bench Verified, HumanEval and MMLU fail on all three counts.

The three groups

A benchmark earns its place by what it holds back, not by how hard it looks. Each group links to its rows below.

15 AI benchmarks ranked by concern load, September 2026Each benchmark occupies one row, ordered most trustworthy first. Four fixed columns show the four concern axes; a Medium rating fills one of the two slots in its column and a High rating fills both, so an empty row means nothing is wrong and a full row means every axis is a problem. SWE-bench Pro ranks first with a concern score of 0 out of 8. Full ratings are in the data table below.Medium concern (1)High concern (2)No concernMost trustworthy firstContaminationSaturationGameabilityReal-world gapConcern1. SWE-bench Pro0 / 82. Terminal-Bench 40 / 83. LiveBench1 / 84. OSWorld 2.01 / 85. tau-bench1 / 86. METR Time Horizon2 / 87. ARC-AGI-33 / 88. FrontierMath4 / 89. ARC-AGI-25 / 810. GPQA Diamond5 / 811. LMArena5 / 812. MMLU-Pro5 / 813. SWE-bench Verified7 / 814. HumanEval8 / 815. MMLU8 / 8
15 AI benchmarks ranked by concern load, September 2026. Each axis is rated Low, Medium or High, where higher means more reason to distrust the headline score. Concern is the sum across all four axes, from 0 to 8.
RankBenchmarkContaminationSaturationGameabilityReal-world gapConcernVerdict
1SWE-bench Pro (real GitHub fixes)LowLowLowLow0 of 8trust
2Terminal-Bench 4 (terminal tasks)LowLowLowLow0 of 8trust
3LiveBench (contamination-limited composite)LowLowLowMedium1 of 8trust
4OSWorld 2.0 (computer use)LowLowMediumLow1 of 8trust
5tau-bench (tool use under policy)MediumLowLowLow1 of 8trust
6METR Time Horizon (task length, not accuracy)LowLowMediumMedium2 of 8trust
7ARC-AGI-3 (interactive reasoning)LowLowMediumHigh3 of 8situational
8FrontierMath (research math)MediumHighLowMedium4 of 8situational
9ARC-AGI-2 (abstract reasoning)LowHighMediumHigh5 of 8situational
10GPQA Diamond (PhD-level science)MediumHighMediumMedium5 of 8situational
11LMArena (human preference)MediumLowHighHigh5 of 8situational
12MMLU-Pro (knowledge MCQ, harder)MediumMediumMediumHigh5 of 8situational
13SWE-bench Verified (real GitHub fixes)HighHighHighMedium7 of 8stop-quoting
14HumanEval (function-level coding)HighHighHighHigh8 of 8stop-quoting
15MMLU (knowledge MCQ)HighHighHighHigh8 of 8stop-quoting
Concern load across the four axes, ranked. Position is the axis, never colour, so the columns read the same whether or not the risk hues separate for you. Ratings are a Capital and Compute synthesis of documented properties, verified September 2026; nothing on this page is hands-on tested.
#BenchmarkContaminationSaturationGameabilityReal-world gapConcernVerdict
1SWE-bench Proreal GitHub fixesLowLowLowLow0 / 8Trust
2Terminal-Bench 4terminal tasksLowLowLowLow0 / 8Trust
3LiveBenchcontamination-limited compositeLowLowLowMedium1 / 8Trust
4OSWorld 2.0computer useLowLowMediumLow1 / 8Trust
5tau-benchtool use under policyMediumLowLowLow1 / 8Trust
6METR Time Horizontask length, not accuracyLowLowMediumMedium2 / 8Trust
7ARC-AGI-3interactive reasoningLowLowMediumHigh3 / 8Caveat
8FrontierMathresearch mathMediumHighLowMedium4 / 8Caveat
9ARC-AGI-2abstract reasoningLowHighMediumHigh5 / 8Caveat
10GPQA DiamondPhD-level scienceMediumHighMediumMedium5 / 8Caveat
11LMArenahuman preferenceMediumLowHighHigh5 / 8Caveat
12MMLU-Proknowledge MCQ, harderMediumMediumMediumHigh5 / 8Caveat
13SWE-bench Verifiedreal GitHub fixesHighHighHighMedium7 / 8Stop
14HumanEvalfunction-level codingHighHighHighHigh8 / 8Stop
15MMLUknowledge MCQHighHighHighHigh8 / 8Stop
Higher is worse on every axis, which is what lets the four sum into one concern score. Names link to the full profile in the benchmarks directory, where the current record score lives. Scores are deliberately not duplicated here: they move, and a second copy would go stale.

The four axes, and why they sum

A benchmark score is a claim about a model. Four things can break that claim, and they break it in different ways, so each is rated separately and then added. All four are scored so that higher always means worse, which is the only reason a single concern number out of 8 is meaningful rather than an average of incompatible things.

  • Contamination. How likely the test data leaked into training. Higher = worse.
  • Saturation. How close top models sit to the ceiling, so the score stops telling models apart. Higher = worse.
  • Gameability. How easily a score is inflated without real capability. Higher = worse.
  • Real-world gap. Gap between the headline score and real task performance. Higher = worse.

The bands are an editorial line, stated so you can disagree with it precisely. A benchmark carrying at most one Medium and nothing High is in the trust group. Anything from three to five is situational: it measures something real and carries one problem large enough that the number needs a sentence beside it. Six or more, which in practice means two or more High ratings, is in the stop-quoting group.

What the axes deliberately do not measure is difficulty. A hard benchmark is not a trustworthy one, and several of the hardest here sit mid-table: ARC-AGI-2 is brutal and nearly contamination-proof, and it still lands in the situational group because the headroom is gone and the thing it measures is fluid reasoning rather than work anyone is paying for.

The 6 benchmarks worth trusting

The pattern in this group is not cleverness. It is what each one refuses to publish, and what it does instead of asking a model to pick from a list. Most of them grade by executing something rather than by comparing text, the rest keep their questions moving, and none of them is close to a ceiling. The top two tie with nothing flagged at all.

1. SWE-bench Pro

real GitHub fixes. Concern 0 of 8. The replacement OpenAI itself now recommends over Verified. Held-out and commercial splits keep the answers off GitHub, and standardized scaffolding removes the harness inflation that lets vendors publish their own numbers. The effect is blunt: models near 80% on Verified land in the 45 to 60% range here, which is the size of the contamination and scaffold premium the older benchmark was hiding. Full SWE-bench Pro profile.

2. Terminal-Bench 4

terminal tasks. Concern 0 of 8. Graded pass/fail by a verifier running in a separate container from the agent, against real test suites in a real terminal, which is hard to fake. The task set is written for the benchmark by a community of 100+ contributors and reviewers rather than scraped from public issues, so there is nothing public to memorize. Saturation is actively managed rather than tolerated: v4.0 deleted tasks every current-generation model solved 5/5. Two real caveats. Versions are not comparable, since v2.1, v3.0 and v4.0 are different task sets sharing a name and differ by more than 30 points. And leakage rises as each public set ages, which is why the operators warn against training on it. Full Terminal-Bench 4 profile.

3. LiveBench

contamination-limited composite. Concern 1 of 8. The strongest composite alternative to preference arenas: no LLM judge, no human votes, and monthly question rotation that limits what can be memorized. Its authors renamed it from contamination-free to contamination-limited, which is the honest framing and the reason it scores well here. Real-world gap is medium because the task mix is academic rather than production work. Full LiveBench profile.

4. OSWorld 2.0

computer use. Concern 1 of 8. Exists because agents reached 83.5% on OSWorld 1.0. Real desktop tasks in a real OS, where a skilled human needs a median of roughly 1.6 hours per task, and the headroom is enormous. Marked medium on gaming for one reason: it reports a partial score alongside full completion, and vendors quote the partial figure from their own harnesses. Full OSWorld 2.0 profile.

5. tau-bench

tool use under policy. Concern 1 of 8. Built to expose unreliability rather than to be topped. Scoring uses pass^k, which requires the same task to succeed across k independent trials, so a lucky single run earns nothing: pass^8 falls far below pass^1 for every model tested. The public task set has been available since 2024, which is the one real contamination exposure here. Full tau-bench profile.

6. METR Time Horizon

task length, not accuracy. Concern 2 of 8. The only entry whose unit is time rather than a percentage, so it cannot saturate the way a capped score does: it reports the task length a model completes at 50% success. The launch paper found a doubling roughly every seven months since 2019. Marked medium on gaming and real-world gap because the fitted horizon is a derived statistic that moves with the task mix and the harness, not a directly observed result. Full METR Time Horizon profile.

The 6 that need a caveat

None of these is junk, and quoting any of them bare is still misleading. Each one has a single dominant problem, and knowing which problem it is tells you what the number can and cannot support.

7. ARC-AGI-3

interactive reasoning. Concern 3 of 8. Where the ARC gap moved. Interactive game environments with no instructions, launched March 2026 with humans at 100% and frontier AI at 0.51%, and still far from solved. Scoring on human-level action efficiency fences brute-force search better than v2 did. The real-world gap is intentional: it measures fluid reasoning, not production work. Harness choice is load-bearing, which is the gaming exposure. Full ARC-AGI-3 profile.

8. FrontierMath

research math. Concern 4 of 8. Held back from the public and from most labs, and designed against guessing, so it resists gaming better than anything else here. The contamination rating is not about leakage to the field: Epoch discloses that OpenAI commissioned and funds it, owns the problems, and has problem and solution access outside a held-out set. That is a conflict of interest for one lab specifically. Now saturated on the graded tiers, leaving the Open Problems track as the frontier. Full FrontierMath profile.

9. ARC-AGI-2

abstract reasoning. Concern 5 of 8. Private and semi-private eval sets keep contamination low, but the headroom is gone. Verified scores went from 54% in December 2025 to past the 85% target within seven months. Gameable by high-compute brute force, fenced by cost-per-task caps rather than by design. The real-world gap is intentional: it measures fluid reasoning, not production work. Full ARC-AGI-2 profile.

10. GPQA Diamond

PhD-level science. Concern 5 of 8. Google-proof PhD science questions with a 65% expert baseline and 34% for skilled non-experts with web access. The reasoning-gated design resists pure retrieval far better than MMLU, but top models are now past the human-expert baseline, so it has largely saturated. Only 198 items, which means a handful of questions swings the headline number. Full GPQA Diamond profile.

11. LMArena

human preference. Concern 5 of 8. Measures preference, not capability, and style is load-bearing: length, formatting and tone move votes. Meta topped it in April 2025 with a chat-tuned variant that was not the released model, and the operator said afterwards that this violated its expectations. The Leaderboard Illusion paper documents how privileged private testing and selective disclosure let large labs overfit the ranking. Full LMArena profile.

12. MMLU-Pro

knowledge MCQ, harder. Concern 5 of 8. Built to replace saturated MMLU: ten answer options instead of four and reasoning-heavy items drop scores 16 to 33 points and separate frontier models again. It is the right substitute if a knowledge score is required, but it is still public multiple choice, and by mid-2026 the top tier is compressing again, so it is walking the same path. Full MMLU-Pro profile.

The 3 benchmarks to stop quoting

These three share a profile: public test data, top scores bunched at the ceiling, and a documented reason the field has already moved on. A record on any of them is better read as evidence of contamination than of capability, and the most striking case is the one that is still quoted most.

13. SWE-bench Verified

real GitHub fixes. Concern 7 of 8. The most-quoted coding number in launch posts, and OpenAI stopped reporting it in February 2026 because frontier models reproduce the reference patches. The gold patches are public on GitHub. SWE-Bench+ found answer leakage and weak tests. Top models sit near the ceiling. Superseded in practice by SWE-bench Pro and by execution-graded successors, and still quoted anyway. Full SWE-bench Verified profile.

14. HumanEval

function-level coding. Concern 8 of 8. Effectively solved: frontier models sit near 99% pass@1, so it cannot separate strong models at all. 164 short standalone functions with public tests, present in essentially every code training corpus. It survives as a historical baseline and as a number to pad a launch table with, not as evidence about engineering ability. Full HumanEval profile.

15. MMLU

knowledge MCQ. Concern 8 of 8. Fully public four-option multiple choice, and demonstrably memorized: GPT-4 reconstructed masked answer options 57% of the time in a peer-reviewed contamination study. Top models bunch in the low-to-mid 90s, so rank differences are mostly noise. The field already built its replacements, MMLU-Pro and MMLU-Redux, and still quotes the original. Full MMLU profile.

What to quote instead

The substitutions are close to one-for-one, which removes the usual excuse for leaning on a saturated number. For coding, SWE-bench Pro replaces SWE-bench Verified and the same models drop from around 80 percent into the 45 to 60 percent range, which is the size of the premium the old benchmark was hiding. For knowledge, MMLU-Pro replaces MMLU and takes 16 to 33 points off with ten answer options instead of four, though it is walking the same road and will need replacing too. For function-level coding, nothing replaces HumanEval, because nothing needs to: it measures whether a model can write a short standalone function, and every frontier model can.

For agentic work the honest answer is that there is no single number. Terminal-Bench 4 covers terminal tasks, tau-bench covers tool use under a policy, and OSWorld 2.0 covers driving a desktop. They measure different jobs and disagree about which model is best, which is information rather than a defect. Thebenchmarks directory carries the current record on each, and the value leaderboard puts capability against what it costs to obtain.

One habit does more than any substitution: quote the version. Terminal-Bench v2.1, v3.0 and v4.0 are different task sets sharing a name, and the top score differs across them by more than 30 points. A benchmark number without a version is not a claim anyone can check, and versionless comparisons are the most common way an honest score becomes a misleading one.

How this ranking was made, and what it cannot tell you

The ratings are judgement, not measurement, and the distinction matters enough to state plainly. No score in the table above was produced by this site, nothing here is hands-on tested, and each of the four axes is scored from evidence that is public and linkable: whether the test set is published, whether frontier scores have bunched, whether the funding arrangement gives one lab an advantage, and whether the harness moves the result. Every row links the primary source its rating rests on, so a disagreement can be argued at the evidence rather than about the conclusion.

Three limits are worth naming. Concern scores are ordinal, so a benchmark at 4 is not twice as compromised as one at 2. The axes are unweighted, and there is a real argument that contamination should count for more than real-world gap on a page about whether to believe a number. And the ranking moves faster than it looks: ARC-AGI-2 went from 54 percent in December 2025 to past its 85 percent target within seven months, which would have moved it a whole band in that time. Current headline scores are read from thebenchmarks directory rather than copied onto this page, so a score that moves cannot leave a contradicting duplicate behind, but the trust ratings themselves need re-checking as the field saturates its own tests.

Frequently asked questions

Which AI benchmark should you trust for coding?
SWE-bench Pro and Terminal-Bench 4, which tie at the top with nothing flagged on any of the four axes. Both execute code and grade pass or fail against real test suites, both keep their answers off public GitHub, and both sit far from their ceiling. Choose between them by the job: Pro for repository-level issue fixing, Terminal-Bench for whole-task work in a shell. SWE-bench Verified is the number you will see quoted instead, and it is the one benchmark here whose most prominent user stopped reporting it.
Why is SWE-bench Verified still quoted if OpenAI stopped reporting it?
Because it is legible and the alternatives are not yet. Verified has years of published scores behind it, so a launch post can show a number that readers recognize and compare. SWE-bench Pro drops the same models from around 80 percent into the 45 to 60 percent range, which is a worse headline for an identical model. The incentive runs toward the older number, not the better one.
What does benchmark contamination actually mean?
That the test questions, or the answers, ended up in the training data, so the model is partly recalling rather than reasoning. It is measurable: a peer-reviewed study found GPT-4 could reconstruct masked MMLU answer options 57 percent of the time. Once a public test set is a few years old, contamination is the default assumption, not the exception.
What does saturation mean for a launch announcement?
That the benchmark has stopped separating models, so a new record on it is not evidence of much. When every frontier model scores in the low-to-mid 90s on MMLU or near 99 percent on HumanEval, the remaining spread is closer to noise than to capability. A saturated benchmark in a launch table is padding.
Can a vendor game a benchmark score?
Yes, and the documented cases are not subtle. Meta topped LMArena in April 2025 with a chat-tuned variant that was not the model it released, and the operator said afterwards that this violated its expectations. More routinely, vendors publish results from their own harness rather than the public board, which is why scaffold-sensitive benchmarks are marked as gameable here even when the task set is clean.
Which AI benchmarks are private or held back?
FrontierMath is held back from the public and from most labs, and the ARC-AGI sets keep semi-private and private splits. Holding data back is the strongest single defence against contamination, and it costs something: you cannot audit what you cannot see, and FrontierMath carries a disclosed funding conflict with one lab that has problem and solution access.
Are these ratings measured or judged?
Judged, from documented properties, and the page says so rather than implying a measurement. Each of the four axes is scored Low, Medium or High on evidence that is public and linkable: whether the test set is published, whether top scores have bunched at the ceiling, who funds the benchmark, and whether the harness moves the result. Nothing here is hands-on tested, and no score in the table was produced by this site.
How current is this ranking?
Every row was re-verified against its primary source in September 2026, and current headline scores are read from the benchmarks directory rather than copied here, so a score that moves cannot leave a stale duplicate behind. Benchmarks saturate fast: ARC-AGI-2 went from 54 percent to past its 85 percent target inside seven months. Check the linked source before quoting any figure.

Sources

Every row above links the primary source its rating rests on. The evidence the four axes lean on hardest:

  • Deng, Zhao, Tang, Gerstein and Cohan (2024). Investigating Data Contamination in Modern Benchmarks for Large Language Models. NAACL 2024 (peer-reviewed).aclanthology.org
  • OpenAI (2026). Why we no longer evaluate SWE-bench Verified. Vendor statement.openai.com
  • Singh et al. (2025). The Leaderboard Illusion. Preprint.arxiv.org/abs/2504.20879
  • Akhtar et al. (2026). When AI Benchmarks Plateau. Preprint.arxiv.org/abs/2602.16763
  • Epoch AI (2025). Clarifying the creation and use of FrontierMath. Benchmark operator statement.epoch.ai
  • Terminal-Bench (2026). Terminal-Bench 3.0. Benchmark operator announcement.tbench.ai
  • Rein et al. (2023). GPQA: A Graduate-Level Google-Proof Q&A Benchmark. Preprint.arxiv.org/abs/2311.12022