MMLU-Pro
MMLU-Pro is the de-saturated replacement for MMLU. It keeps the multiple-choice format but expands each question from four options to ten and selects harder, reasoning-heavy items, which lowers scores enough to separate frontier models again. By mid-2026 the top of the range is compressing, so its useful life is also finite.
| What it measures | Harder multi-task reasoning and knowledge designed to de-saturate MMLU and reward deliberate reasoning over recall. |
|---|---|
| Built by | TIGER-Lab (Wang et al., University of Waterloo), 2024 |
| Format | About 12,000 questions across 14 disciplines, expanded from 4 to 10 answer options to cut the guessing baseline |
| Scoring metric | Accuracy |
| Status | Active |
| Representative top score | ~90% · Gemini 3 Pro Preview · read 2026-06 |
| Official leaderboard | huggingface.co/spaces/TIGER-Lab/MMLU-Pro |
How MMLU-Pro works
The set contains roughly 12,000 questions across 14 disciplines. The move from four options to ten cuts the random-guess baseline from 25% to 10%, which removes a large chunk of the score a weak model could previously obtain for free. Item selection favours questions that require working through a problem rather than recalling a fact, so chain-of-thought prompting helps materially more here than on MMLU.
History and current status
TIGER-Lab at the University of Waterloo released it in 2024 in direct response to MMLU saturation. Reported drops of 16 to 33 points relative to MMLU confirmed it had restored headroom. Through 2026 frontier scores climbed to roughly 90%, and the leading models are again bunching, which puts MMLU-Pro on the same trajectory as its predecessor, just a couple of years behind.
What the score does not tell you
It inherits MMLU’s basic weakness: it is a public, static, multiple-choice set, so contamination accumulates with every training run. Ten options raise the difficulty but do not change the format, and multiple choice rewards elimination strategies that do not correspond to understanding. Its saturation curve shows that harder questions in the same format buy time rather than solving the problem.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Scored on those four axes, MMLU-Pro carries a concern score of 5 out of 8 and ranks 12 of 15, which puts it in the group that needs a caveat beside the number. See the reasoning behind that rating and how it compares with the rest of the field.
What MMLU-Pro scores actually mean
MMLU-Pro was built to restore the resolution MMLU lost, and it worked for about two years. Expanding from four answer options to ten drops the random baseline from 25% to 10%, and the reasoning-heavy item selection cut frontier scores by 16 to 33 points relative to MMLU at release. Today the leaders read around 90%, which means the same compression is returning: the top tier is bunching, and the distance between the best and fifth-best model is again small enough to be scaffolding rather than capability. In the 60% to 85% band the score still separates models meaningfully, which is where most open-weight and mid-tier commercial models sit.
Who reports MMLU-Pro, and how to read it
MMLU-Pro is now the standard broad-knowledge row in frontier model cards, having displaced MMLU in most launch tables. Because scores are converging near the top, read it alongside an unsaturated reasoning benchmark rather than as the headline capability number. It retains real resolution for open-weight and mid-tier models, which is where it is still worth quoting.
When to weight MMLU-Pro in a model choice
Use MMLU-Pro when comparing open-weight or mid-tier models, where it still has real resolution, and discount it when comparing frontier models, where it no longer does. It is a knowledge-and-reasoning breadth test, so it predicts general question answering far better than it predicts agentic or coding performance; do not let a strong MMLU-Pro number stand in for either. If the models under comparison all read above 88%, the benchmark has run out of signal for that comparison and you should move to GPQA Diamond, Humanity's Last Exam or a task-specific evaluation instead.
Benchmarks to read alongside this one
MMLU
Broad academic and professional knowledge across 57 subjects via four-choice multiple-choice questions.
MMLU-Redux
A re-annotated, error-corrected subset of MMLU used to measure true knowledge accuracy without the original's label noise.
SuperGPQA
Graduate-level knowledge and reasoning across 285 disciplines, including the applied and service fields that mainstream benchmarks ignore.
MMLU-Pro: frequently asked questions
- What is MMLU-Pro?
- MMLU-Pro is a 2024 benchmark of about 12,000 questions across 14 disciplines, built to replace the saturated MMLU. It expands each question from four answer options to ten and favours reasoning-heavy items, which lowers scores and restores separation between strong models.
- Why is MMLU-Pro harder than MMLU?
- Two reasons. Ten answer options instead of four cut the random-guess baseline from 25% to 10%, and the questions were selected to require multi-step reasoning rather than recall. Together these drop model scores by roughly 16 to 33 points.
- Is MMLU-Pro saturated?
- Not yet, but it is heading that way. Frontier models reached roughly 90% by mid-2026 and the leaders are compressing into a narrow band, which is the same pattern MMLU showed before it stopped being useful.
Sources
- Wang, Ma, Zhang, Ni et al. (2024). MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. arXiv preprint 2406.01574. arxiv.org/abs/2406.01574
- TIGER-Lab. MMLU-Pro leaderboard on Hugging Face Spaces. huggingface.co/spaces/TIGER-Lab/MMLU-Pro
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.
← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →