Capital & Compute
Benchmark· Knowledge & general QA· Checked 2026-06-29

MMLU

Also known as Massive Multitask Language Understanding

MMLU, or Massive Multitask Language Understanding, is a multiple-choice exam covering 57 academic and professional subjects. For years it was the single most-quoted number in AI, the default proxy for "how smart is this model." In 2026 it is saturated, demonstrably contaminated, and known to contain wrong answers, so a high MMLU score carries almost no information.

Key facts about the MMLU benchmark
What it measuresBroad academic and professional knowledge across 57 subjects via four-choice multiple-choice questions.
Built byHendrycks et al. (UC Berkeley and collaborators), 2021
FormatAbout 15,900 four-option questions across 57 subjects (STEM, humanities, social sciences, professional exams)
Scoring metricAccuracy
StatusSaturated
Representative top score~93% · Qwen3.7 Max · read 2026-06
Official leaderboardllm-stats.com/benchmarks/mmlu

How MMLU works

The benchmark presents roughly 15,900 four-option multiple-choice questions spanning STEM, the humanities, the social sciences and professional exams such as law and medicine. Scoring is plain accuracy, with a random-guess baseline of 25%. Because it is multiple choice and fully automated, it is cheap to run, which is a large part of why it became ubiquitous.

History and current status

Introduced in a 2021 paper by Hendrycks and collaborators, when strong models scored barely above chance on many subjects, MMLU became the headline row in every model card through the GPT-3 and GPT-4 era. Frontier models passed the 90% mark, and by mid-2026 the top of the range sits around 93%. Its successors, MMLU-Pro and MMLU-Redux, exist specifically because it stopped working.

What the score does not tell you

MMLU has three separate problems. It is saturated, so the remaining spread between frontier models is mostly noise. It is contaminated, because the questions have circulated online for years. And independent analysis found genuine errors in its ground-truth answers, meaning a perfect model could not score 100%. MMLU-Redux was built to correct that last problem specifically.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, why benchmark saturation makes a high score meaningless.

Who reports MMLU, and how to read it

Model developers still include MMLU in comparison tables, largely out of convention and because the number is high. Treat its appearance as a signal about the table rather than the model: if a launch leads with MMLU rather than a current unsaturated benchmark, that is a choice about presentation.

2021
First released
Hendrycks et al. (UC Berkeley and collaborators)
Saturated
Status today
As of July 27, 2026
~93%
Representative top score
Read 2026-06

Benchmarks to read alongside this one

MMLU: frequently asked questions

What is the MMLU benchmark?
MMLU is a benchmark of about 15,900 four-option multiple-choice questions across 57 subjects, from elementary mathematics to professional law and medicine. It measures broad academic and professional knowledge by plain accuracy, with a 25% random baseline.
Is MMLU still relevant in 2026?
No, not as a way to rank frontier models. All leading models score above 90%, the questions have leaked into training data, and the answer key itself contains errors. MMLU-Pro, MMLU-Redux, GPQA Diamond and Humanity’s Last Exam are the working replacements.
What is a good MMLU score?
In 2026 anything below roughly 85% marks a model as behind the frontier, and everything above that is compressed into a band where differences are not meaningful. That compression is precisely what makes the benchmark unusable for ranking.
What is the difference between MMLU and MMLU-Pro?
MMLU-Pro raises the number of answer options from four to ten, which cuts the guessing baseline, and selects harder reasoning-heavy questions. That drops scores by roughly 16 to 33 points and restores some separation between frontier models.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory