Capital & Compute

Humanity's Last Exam

Benchmark· Reasoning & abstraction· Checked 2026-07-26

Also known as HLE

Humanity’s Last Exam is a closed-ended examination of expert knowledge and reasoning across more than 100 academic disciplines, written to sit at the limit of what human specialists can answer. Models scored single digits at its release in early 2025. It remains the hardest unsaturated knowledge benchmark in this directory, with the top score around 53%.

Key facts about the Humanity's Last Exam benchmark
What it measuresFrontier, closed-ended expert knowledge and reasoning across more than 100 academic disciplines at the limit of human expertise.
Built byCenter for AI Safety (CAIS) and Scale AI, 2025
Format2,500 public expert-level questions (text and multimodal) across 100+ subjects, mostly short-answer and multiple-choice, plus a private held-out set
Scoring metricAccuracy (exact match / multiple-choice), often reported with a calibration metric
StatusActive
Representative top score53.3% · Claude Fable 5 (Max Effort) · read 2026-06
Official leaderboardartificialanalysis.ai/evaluations/humanitys-last-exam

How Humanity's Last Exam works

The public set contains 2,500 expert-level questions across over 100 subjects, both text and multimodal, mostly short answer and multiple choice. There is also a private held-out set, which exists to detect models that have overfitted to the public questions. Scoring is accuracy, usually reported with a calibration measure so a model that guesses confidently is distinguishable from one that knows.

History and current status

The Center for AI Safety and Scale AI released it in January 2025, and the work was published in Nature in January 2026. Progress has been substantial but far from complete: the recorded top is 53.3% for Claude Fable 5 at maximum effort as of June 2026, with Claude Opus 5 second at 52.6% after a July recheck. The private holdout has so far not revealed large overfitting gaps.

What the score does not tell you

The headline number is easy to misread because conditions differ. Anthropic reported 57.4% for Claude Sonnet 5, but that is a with-tools result rather than the closed-book condition this directory tracks, so the two are not comparable. More broadly, an expert-trivia exam measures retrieval-plus-reasoning breadth, which is not obviously the capability that matters for work, and it is co-run by a vendor.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

Humanity's Last Exam is not one of the 15 benchmarks scored in the benchmark trust ranking, which rates the most-quoted tests on those same four axes and sorts them most trustworthy first.

What Humanity's Last Exam scores actually mean

HLE is one of the few entries here where the numbers still have room to move. Early-2025 models scored in the single digits; the leader now reads 53.3%, so roughly half of 2,500 expert-written questions across 100-plus subjects remain unsolved. That makes a ten-point gap on HLE a real capability difference rather than harness noise, which is the opposite of the situation on MMLU or GPQA. The critical caveat is the evaluation condition: closed-book and tool-assisted runs are different benchmarks wearing one name, and a with-tools figure can sit several points above a closed-book one from the same model. Calibration is often reported alongside accuracy, and a confidently wrong model is worse than an uncertain one.

Who reports Humanity's Last Exam, and how to read it

Frontier labs quote it frequently because it is one of the few knowledge benchmarks with headroom left. Artificial Analysis maintains an independent evaluation, which is the number to prefer. Always check whether a quoted figure is closed-book or tool-assisted, and what effort setting produced it. Those two qualifiers move the number more than most model-to-model differences do.

When to weight Humanity's Last Exam in a model choice

Weight HLE when the use case is hard, multi-domain expert reasoning and you need a benchmark that still separates the frontier. Always record whether the figure is closed-book or with tools before comparing two models, because that single distinction accounts for more spread than most model-to-model differences. Read the calibration number alongside the accuracy one if it is published. HLE says little about coding, tool use or long-horizon agentic reliability, so pair it with SWE-bench Pro, BFCL or Terminal-Bench rather than treating a strong HLE score as general competence.

2025
First released
Center for AI Safety (CAIS) and Scale AI
Active
Status today
As of September 12, 2026
53.3%
Representative top score
Read 2026-06

Benchmarks to read alongside this one

Humanity's Last Exam: frequently asked questions

What is Humanity’s Last Exam?
It is a benchmark of 2,500 public expert-level questions across more than 100 academic disciplines, released in January 2025 by the Center for AI Safety and Scale AI, with a private held-out set. It measures frontier closed-ended expert knowledge and reasoning, scored by accuracy.
What is the top score on Humanity’s Last Exam?
About 53% in the closed-book condition this directory tracks, held by Claude Fable 5 at maximum effort as of June 2026. Higher figures circulate, but they generally come from tool-assisted runs, which are not comparable to closed-book scores.
Why is Humanity’s Last Exam not saturated?
Because the questions were written by specialists to sit at the edge of human expertise across a very wide range of fields, so there is no shortcut through breadth. Models went from single digits in early 2025 to roughly 53% by mid-2026, which leaves considerable headroom.

Sources

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →