MMMU
Also known as Massive Multi-discipline Multimodal Understanding
MMMU is the standard expert-level test of multimodal understanding: college-level questions requiring reasoning over images, diagrams, charts and text together. It is the benchmark most often cited when a model claims to see, and it also has a well-documented flaw, which is that a meaningful share of its questions can be answered without looking at the image at all.
| What it measures | College-level multimodal understanding and reasoning over images, diagrams, charts and text across many disciplines. |
|---|---|
| Built by | MMMU team (Yue et al.), 2023 |
| Format | About 11,500 questions across 6 disciplines and 30 subjects, mixing multiple-choice and open-ended items with 30+ image types |
| Scoring metric | Accuracy |
| Status | Active |
| Representative top score | ~86% · Qwen3.6 Plus · read 2026-06 |
| Official leaderboard | mmmu-benchmark.github.io |
How MMMU works
The set contains roughly 11,500 questions across six disciplines and 30 subjects, mixing multiple-choice and open-ended items, with more than 30 distinct image types including diagrams, charts, tables, chemical structures and musical notation. Scoring is accuracy. Because the questions are drawn from college-level material, the intended bar is expert rather than general competence.
History and current status
The MMMU team introduced it in 2023, when GPT-4V scored roughly 56%, leaving substantial headroom. Frontier models have since closed much of that gap, with the top around 86% by mid-2026. MMMU-Pro was released as the harder variant, and the same period produced MMStar, which was built specifically to measure how much of an MMMU-style score is genuinely visual.
What the score does not tell you
The MMStar authors documented the central problem: a text-only model scored 42.9% on MMMU with no image input at all, and beat the random baseline by over 24% on average across six vision benchmarks. That means a large part of an MMMU score reflects language priors and answer-option elimination rather than seeing. A high MMMU number is therefore weak evidence of visual reasoning on its own.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Who reports MMMU, and how to read it
MMMU is the default multimodal row in frontier model cards. Read it alongside MMStar, which controls for visual dependency, and CharXiv, which uses real scientific figures rather than clean template charts. Any of those three alone overstates what a vision model can actually do.
Benchmarks to read alongside this one
MMStar
Genuinely vision-dependent multimodal ability, on samples selected so the answer cannot be inferred from the text alone.
MMMU-Pro
A harder, contamination-resistant version of MMMU that forces genuine visual reasoning rather than text-only shortcuts.
CharXiv
Chart understanding on real, messy scientific figures rather than clean template-generated charts.
MMMU: frequently asked questions
- What is the MMMU benchmark?
- MMMU is a benchmark of about 11,500 college-level questions across six disciplines and 30 subjects, requiring reasoning over more than 30 types of image alongside text. It mixes multiple-choice and open-ended items and is scored by accuracy.
- Can models score well on MMMU without seeing the image?
- Partly, yes. The MMStar authors reported a text-only model reaching 42.9% on MMMU with no visual input, which means language priors and option elimination contribute substantially to the score. That is why MMMU should be read alongside a visual-dependency-controlled benchmark.
- What is the difference between MMMU and MMMU-Pro?
- MMMU-Pro is the harder variant, built after frontier models closed much of the headroom on the original. It tightens the evaluation to reduce the share of questions answerable from text alone and to restore separation between strong vision models.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.