MMMU
Also known as Massive Multi-discipline Multimodal Understanding
MMMU is the standard expert-level test of multimodal understanding: college-level questions requiring reasoning over images, diagrams, charts and text together. It is the benchmark most often cited when a model claims to see, and it also has a well-documented flaw, which is that a meaningful share of its questions can be answered without looking at the image at all.
| What it measures | College-level multimodal understanding and reasoning over images, diagrams, charts and text across many disciplines. |
|---|---|
| Built by | MMMU team (Yue et al.), 2023 |
| Format | About 11,500 questions across 6 disciplines and 30 subjects, mixing multiple-choice and open-ended items with 30+ image types |
| Scoring metric | Accuracy |
| Status | Active |
| Representative top score | ~86% · Qwen3.6 Plus · read 2026-06 |
| Official leaderboard | mmmu-benchmark.github.io |
How MMMU works
The set contains roughly 11,500 questions across six disciplines and 30 subjects, mixing multiple-choice and open-ended items, with more than 30 distinct image types including diagrams, charts, tables, chemical structures and musical notation. Scoring is accuracy. Because the questions are drawn from college-level material, the intended bar is expert rather than general competence.
History and current status
The MMMU team introduced it in 2023, when GPT-4V scored roughly 56%, leaving substantial headroom. Frontier models have since closed much of that gap, with the top around 86% by mid-2026. MMMU-Pro was released as the harder variant, and the same period produced MMStar, which was built specifically to measure how much of an MMMU-style score is genuinely visual.
What the score does not tell you
The MMStar authors documented the central problem: a text-only model scored 42.9% on MMMU with no image input at all, and beat the random baseline by over 24% on average across six vision benchmarks. That means a large part of an MMMU score reflects language priors and answer-option elimination rather than seeing. A high MMMU number is therefore weak evidence of visual reasoning on its own.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
What MMMU scores actually mean
MMMU has moved from a wide-open benchmark to a narrowing one. GPT-4V scored about 56% at release in 2023; the leaders now read around 86% across roughly 11,500 questions spanning six disciplines and 30-plus image types. Most of the remaining headroom sits in the image-heavy disciplines where the question cannot be answered from the text alone, and a large share of items can be partially inferred without truly reading the figure. That makes the aggregate number a generous estimate of visual reasoning: models that score well can be weak on diagram interpretation specifically, and the per-discipline breakdown is where that shows up.
Who reports MMMU, and how to read it
MMMU is the default multimodal row in frontier model cards. Read it alongside MMStar, which controls for visual dependency, and CharXiv, which uses real scientific figures rather than clean template charts. Any of those three alone overstates what a vision model can actually do. The per-discipline breakdown varies more than the gap between frontier models, so the aggregate hides the answer you need.
When to weight MMMU in a model choice
Use MMMU as the general check on whether a model can reason over expert-level material that includes charts, diagrams and technical images, which is the common failure mode when a text-strong model is put in front of documents. Read the per-discipline results rather than the aggregate, especially for the domain you actually care about, since the spread across disciplines is wider than the spread across frontier models. Do not treat MMMU as a proxy for document extraction, OCR quality or chart-to-data accuracy; those are separate capabilities that a strong MMMU score does not guarantee.
Benchmarks to read alongside this one
MMStar
Genuinely vision-dependent multimodal ability, on samples selected so the answer cannot be inferred from the text alone.
MMMU-Pro
A harder, contamination-resistant version of MMMU that forces genuine visual reasoning rather than text-only shortcuts.
CharXiv
Chart understanding on real, messy scientific figures rather than clean template-generated charts.
MMMU: frequently asked questions
- What is the MMMU benchmark?
- MMMU is a benchmark of about 11,500 college-level questions across six disciplines and 30 subjects, requiring reasoning over more than 30 types of image alongside text. It mixes multiple-choice and open-ended items and is scored by accuracy.
- Can models score well on MMMU without seeing the image?
- Partly, yes. The MMStar authors reported a text-only model reaching 42.9% on MMMU with no visual input, which means language priors and option elimination contribute substantially to the score. That is why MMMU should be read alongside a visual-dependency-controlled benchmark.
- What is the difference between MMMU and MMMU-Pro?
- MMMU-Pro is the harder variant, built after frontier models closed much of the headroom on the original. It tightens the evaluation to reduce the share of questions answerable from text alone and to restore separation between strong vision models.
Sources
- Yue, Ni, Zhang, Zheng et al. (2023). MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI. arXiv preprint 2311.16502. arxiv.org/abs/2311.16502
- MMMU team. Official benchmark site and leaderboard. mmmu-benchmark.github.io
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.