MMLU-Pro
MMLU-Pro is the de-saturated replacement for MMLU. It keeps the multiple-choice format but expands each question from four options to ten and selects harder, reasoning-heavy items, which lowers scores enough to separate frontier models again. By mid-2026 the top of the range is compressing, so its useful life is also finite.
| What it measures | Harder multi-task reasoning and knowledge designed to de-saturate MMLU and reward deliberate reasoning over recall. |
|---|---|
| Built by | TIGER-Lab (Wang et al., University of Waterloo), 2024 |
| Format | About 12,000 questions across 14 disciplines, expanded from 4 to 10 answer options to cut the guessing baseline |
| Scoring metric | Accuracy |
| Status | Active |
| Representative top score | ~90% · Gemini 3 Pro Preview · read 2026-06 |
| Official leaderboard | huggingface.co/spaces/TIGER-Lab/MMLU-Pro |
How MMLU-Pro works
The set contains roughly 12,000 questions across 14 disciplines. The move from four options to ten cuts the random-guess baseline from 25% to 10%, which removes a large chunk of the score a weak model could previously obtain for free. Item selection favours questions that require working through a problem rather than recalling a fact, so chain-of-thought prompting helps materially more here than on MMLU.
History and current status
TIGER-Lab at the University of Waterloo released it in 2024 in direct response to MMLU saturation. Reported drops of 16 to 33 points relative to MMLU confirmed it had restored headroom. Through 2026 frontier scores climbed to roughly 90%, and the leading models are again bunching, which puts MMLU-Pro on the same trajectory as its predecessor, just a couple of years behind.
What the score does not tell you
It inherits MMLU’s basic weakness: it is a public, static, multiple-choice set, so contamination accumulates with every training run. Ten options raise the difficulty but do not change the format, and multiple choice rewards elimination strategies that do not correspond to understanding. Its saturation curve shows that harder questions in the same format buy time rather than solving the problem.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Who reports MMLU-Pro, and how to read it
MMLU-Pro is now the standard broad-knowledge row in frontier model cards, having displaced MMLU in most launch tables. Because scores are converging near the top, read it alongside an unsaturated reasoning benchmark rather than as the headline capability number.
Benchmarks to read alongside this one
MMLU
Broad academic and professional knowledge across 57 subjects via four-choice multiple-choice questions.
MMLU-Redux
A re-annotated, error-corrected subset of MMLU used to measure true knowledge accuracy without the original's label noise.
SuperGPQA
Graduate-level knowledge and reasoning across 285 disciplines, including the applied and service fields that mainstream benchmarks ignore.
MMLU-Pro: frequently asked questions
- What is MMLU-Pro?
- MMLU-Pro is a 2024 benchmark of about 12,000 questions across 14 disciplines, built to replace the saturated MMLU. It expands each question from four answer options to ten and favours reasoning-heavy items, which lowers scores and restores separation between strong models.
- Why is MMLU-Pro harder than MMLU?
- Two reasons. Ten answer options instead of four cut the random-guess baseline from 25% to 10%, and the questions were selected to require multi-step reasoning rather than recall. Together these drop model scores by roughly 16 to 33 points.
- Is MMLU-Pro saturated?
- Not yet, but it is heading that way. Frontier models reached roughly 90% by mid-2026 and the leaders are compressing into a narrow band, which is the same pattern MMLU showed before it stopped being useful.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.