Humanity's Last Exam
Also known as HLE
Humanity’s Last Exam is a closed-ended examination of expert knowledge and reasoning across more than 100 academic disciplines, written to sit at the limit of what human specialists can answer. Models scored single digits at its release in early 2025. It remains the hardest unsaturated knowledge benchmark in this directory, with the top score around 53%.
| What it measures | Frontier, closed-ended expert knowledge and reasoning across more than 100 academic disciplines at the limit of human expertise. |
|---|---|
| Built by | Center for AI Safety (CAIS) and Scale AI, 2025 |
| Format | 2,500 public expert-level questions (text and multimodal) across 100+ subjects, mostly short-answer and multiple-choice, plus a private held-out set |
| Scoring metric | Accuracy (exact match / multiple-choice), often reported with a calibration metric |
| Status | Active |
| Representative top score | 53.3% · Claude Fable 5 (Max Effort) · read 2026-06 |
| Official leaderboard | artificialanalysis.ai/evaluations/humanitys-last-exam |
How Humanity's Last Exam works
The public set contains 2,500 expert-level questions across over 100 subjects, both text and multimodal, mostly short answer and multiple choice. There is also a private held-out set, which exists to detect models that have overfitted to the public questions. Scoring is accuracy, usually reported with a calibration measure so a model that guesses confidently is distinguishable from one that knows.
History and current status
The Center for AI Safety and Scale AI released it in January 2025, and the work was published in Nature in January 2026. Progress has been substantial but far from complete: the recorded top is 53.3% for Claude Fable 5 at maximum effort as of June 2026, with Claude Opus 5 second at 52.6% after a July recheck. The private holdout has so far not revealed large overfitting gaps.
What the score does not tell you
The headline number is easy to misread because conditions differ. Anthropic reported 57.4% for Claude Sonnet 5, but that is a with-tools result rather than the closed-book condition this directory tracks, so the two are not comparable. More broadly, an expert-trivia exam measures retrieval-plus-reasoning breadth, which is not obviously the capability that matters for work, and it is co-run by a vendor.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Humanity's Last Exam is not one of the 15 benchmarks scored in the benchmark trust ranking, which rates the most-quoted tests on those same four axes and sorts them most trustworthy first.
What Humanity's Last Exam scores actually mean
HLE is one of the few entries here where the numbers still have room to move. Early-2025 models scored in the single digits; the leader now reads 53.3%, so roughly half of 2,500 expert-written questions across 100-plus subjects remain unsolved. That makes a ten-point gap on HLE a real capability difference rather than harness noise, which is the opposite of the situation on MMLU or GPQA. The critical caveat is the evaluation condition: closed-book and tool-assisted runs are different benchmarks wearing one name, and a with-tools figure can sit several points above a closed-book one from the same model. Calibration is often reported alongside accuracy, and a confidently wrong model is worse than an uncertain one.
Who reports Humanity's Last Exam, and how to read it
Frontier labs quote it frequently because it is one of the few knowledge benchmarks with headroom left. Artificial Analysis maintains an independent evaluation, which is the number to prefer. Always check whether a quoted figure is closed-book or tool-assisted, and what effort setting produced it. Those two qualifiers move the number more than most model-to-model differences do.
When to weight Humanity's Last Exam in a model choice
Weight HLE when the use case is hard, multi-domain expert reasoning and you need a benchmark that still separates the frontier. Always record whether the figure is closed-book or with tools before comparing two models, because that single distinction accounts for more spread than most model-to-model differences. Read the calibration number alongside the accuracy one if it is published. HLE says little about coding, tool use or long-horizon agentic reliability, so pair it with SWE-bench Pro, BFCL or Terminal-Bench rather than treating a strong HLE score as general competence.
Benchmarks to read alongside this one
GPQA Diamond
Graduate and PhD-level multiple-choice scientific reasoning in biology, physics and chemistry, on questions designed to be unanswerable by quick web search.
ARC-AGI-3
Whether an agent dropped into an unfamiliar interactive environment with no instructions, stated goal or rules can work out what to do by acting, build a usable world model, and keep learning across levels.
SuperGPQA
Graduate-level knowledge and reasoning across 285 disciplines, including the applied and service fields that mainstream benchmarks ignore.
Humanity's Last Exam: frequently asked questions
- What is Humanity’s Last Exam?
- It is a benchmark of 2,500 public expert-level questions across more than 100 academic disciplines, released in January 2025 by the Center for AI Safety and Scale AI, with a private held-out set. It measures frontier closed-ended expert knowledge and reasoning, scored by accuracy.
- What is the top score on Humanity’s Last Exam?
- About 53% in the closed-book condition this directory tracks, held by Claude Fable 5 at maximum effort as of June 2026. Higher figures circulate, but they generally come from tool-assisted runs, which are not comparable to closed-book scores.
- Why is Humanity’s Last Exam not saturated?
- Because the questions were written by specialists to sit at the edge of human expertise across a very wide range of fields, so there is no shortcut through breadth. Models went from single digits in early 2025 to roughly 53% by mid-2026, which leaves considerable headroom.
Sources
- Phan, Gatti, Han, Li et al. (2025). Humanity's Last Exam. arXiv preprint 2501.14249, later published in Nature (January 2026). arxiv.org/abs/2501.14249
- Artificial Analysis. Independent Humanity's Last Exam evaluation tracker. artificialanalysis.ai/evaluations/humanitys-last-exam
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.
← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →