Capital & Compute

MMLU-Pro

Benchmark· Knowledge & general QA· Checked 2026-06-29

MMLU-Pro is the de-saturated replacement for MMLU. It keeps the multiple-choice format but expands each question from four options to ten and selects harder, reasoning-heavy items, which lowers scores enough to separate frontier models again. By mid-2026 the top of the range is compressing, so its useful life is also finite.

Key facts about the MMLU-Pro benchmark
What it measuresHarder multi-task reasoning and knowledge designed to de-saturate MMLU and reward deliberate reasoning over recall.
Built byTIGER-Lab (Wang et al., University of Waterloo), 2024
FormatAbout 12,000 questions across 14 disciplines, expanded from 4 to 10 answer options to cut the guessing baseline
Scoring metricAccuracy
StatusActive
Representative top score~90% · Gemini 3 Pro Preview · read 2026-06
Official leaderboardhuggingface.co/spaces/TIGER-Lab/MMLU-Pro

How MMLU-Pro works

The set contains roughly 12,000 questions across 14 disciplines. The move from four options to ten cuts the random-guess baseline from 25% to 10%, which removes a large chunk of the score a weak model could previously obtain for free. Item selection favours questions that require working through a problem rather than recalling a fact, so chain-of-thought prompting helps materially more here than on MMLU.

History and current status

TIGER-Lab at the University of Waterloo released it in 2024 in direct response to MMLU saturation. Reported drops of 16 to 33 points relative to MMLU confirmed it had restored headroom. Through 2026 frontier scores climbed to roughly 90%, and the leading models are again bunching, which puts MMLU-Pro on the same trajectory as its predecessor, just a couple of years behind.

What the score does not tell you

It inherits MMLU’s basic weakness: it is a public, static, multiple-choice set, so contamination accumulates with every training run. Ten options raise the difficulty but do not change the format, and multiple choice rewards elimination strategies that do not correspond to understanding. Its saturation curve shows that harder questions in the same format buy time rather than solving the problem.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

Scored on those four axes, MMLU-Pro carries a concern score of 5 out of 8 and ranks 12 of 15, which puts it in the group that needs a caveat beside the number. See the reasoning behind that rating and how it compares with the rest of the field.

What MMLU-Pro scores actually mean

MMLU-Pro was built to restore the resolution MMLU lost, and it worked for about two years. Expanding from four answer options to ten drops the random baseline from 25% to 10%, and the reasoning-heavy item selection cut frontier scores by 16 to 33 points relative to MMLU at release. Today the leaders read around 90%, which means the same compression is returning: the top tier is bunching, and the distance between the best and fifth-best model is again small enough to be scaffolding rather than capability. In the 60% to 85% band the score still separates models meaningfully, which is where most open-weight and mid-tier commercial models sit.

Who reports MMLU-Pro, and how to read it

MMLU-Pro is now the standard broad-knowledge row in frontier model cards, having displaced MMLU in most launch tables. Because scores are converging near the top, read it alongside an unsaturated reasoning benchmark rather than as the headline capability number. It retains real resolution for open-weight and mid-tier models, which is where it is still worth quoting.

When to weight MMLU-Pro in a model choice

Use MMLU-Pro when comparing open-weight or mid-tier models, where it still has real resolution, and discount it when comparing frontier models, where it no longer does. It is a knowledge-and-reasoning breadth test, so it predicts general question answering far better than it predicts agentic or coding performance; do not let a strong MMLU-Pro number stand in for either. If the models under comparison all read above 88%, the benchmark has run out of signal for that comparison and you should move to GPQA Diamond, Humanity's Last Exam or a task-specific evaluation instead.

2024
First released
TIGER-Lab (Wang et al., University of Waterloo)
Active
Status today
As of September 12, 2026
~90%
Representative top score
Read 2026-06

Benchmarks to read alongside this one

MMLU-Pro: frequently asked questions

What is MMLU-Pro?
MMLU-Pro is a 2024 benchmark of about 12,000 questions across 14 disciplines, built to replace the saturated MMLU. It expands each question from four answer options to ten and favours reasoning-heavy items, which lowers scores and restores separation between strong models.
Why is MMLU-Pro harder than MMLU?
Two reasons. Ten answer options instead of four cut the random-guess baseline from 25% to 10%, and the questions were selected to require multi-step reasoning rather than recall. Together these drop model scores by roughly 16 to 33 points.
Is MMLU-Pro saturated?
Not yet, but it is heading that way. Frontier models reached roughly 90% by mid-2026 and the leaders are compressing into a narrow band, which is the same pattern MMLU showed before it stopped being useful.

Sources

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →