Capital & Compute
Benchmark· Reasoning & abstraction· Checked 2026-07-26

Humanity's Last Exam

Also known as HLE

Humanity’s Last Exam is a closed-ended examination of expert knowledge and reasoning across more than 100 academic disciplines, written to sit at the limit of what human specialists can answer. Models scored single digits at its release in early 2025. It remains the hardest unsaturated knowledge benchmark in this directory, with the top score around 53%.

Key facts about the Humanity's Last Exam benchmark
What it measuresFrontier, closed-ended expert knowledge and reasoning across more than 100 academic disciplines at the limit of human expertise.
Built byCenter for AI Safety (CAIS) and Scale AI, 2025
Format2,500 public expert-level questions (text and multimodal) across 100+ subjects, mostly short-answer and multiple-choice, plus a private held-out set
Scoring metricAccuracy (exact match / multiple-choice), often reported with a calibration metric
StatusActive
Representative top score53.3% · Claude Fable 5 (Max Effort) · read 2026-06
Official leaderboardartificialanalysis.ai/evaluations/humanitys-last-exam

How Humanity's Last Exam works

The public set contains 2,500 expert-level questions across over 100 subjects, both text and multimodal, mostly short answer and multiple choice. There is also a private held-out set, which exists to detect models that have overfitted to the public questions. Scoring is accuracy, usually reported with a calibration measure so a model that guesses confidently is distinguishable from one that knows.

History and current status

The Center for AI Safety and Scale AI released it in January 2025, and the work was published in Nature in January 2026. Progress has been substantial but far from complete: the recorded top is 53.3% for Claude Fable 5 at maximum effort as of June 2026, with Claude Opus 5 second at 52.6% after a July recheck. The private holdout has so far not revealed large overfitting gaps.

What the score does not tell you

The headline number is easy to misread because conditions differ. Anthropic reported 57.4% for Claude Sonnet 5, but that is a with-tools result rather than the closed-book condition this directory tracks, so the two are not comparable. More broadly, an expert-trivia exam measures retrieval-plus-reasoning breadth, which is not obviously the capability that matters for work, and it is co-run by a vendor.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

Who reports Humanity's Last Exam, and how to read it

Frontier labs quote it frequently because it is one of the few knowledge benchmarks with headroom left. Artificial Analysis maintains an independent evaluation, which is the number to prefer. Always check whether a quoted figure is closed-book or tool-assisted, and what effort setting produced it.

2025
First released
Center for AI Safety (CAIS) and Scale AI
Active
Status today
As of July 27, 2026
53.3%
Representative top score
Read 2026-06

Benchmarks to read alongside this one

Humanity's Last Exam: frequently asked questions

What is Humanity’s Last Exam?
It is a benchmark of 2,500 public expert-level questions across more than 100 academic disciplines, released in January 2025 by the Center for AI Safety and Scale AI, with a private held-out set. It measures frontier closed-ended expert knowledge and reasoning, scored by accuracy.
What is the top score on Humanity’s Last Exam?
About 53% in the closed-book condition this directory tracks, held by Claude Fable 5 at maximum effort as of June 2026. Higher figures circulate, but they generally come from tool-assisted runs, which are not comparable to closed-book scores.
Why is Humanity’s Last Exam not saturated?
Because the questions were written by specialists to sit at the edge of human expertise across a very wide range of fields, so there is no shortcut through breadth. Models went from single digits in early 2025 to roughly 53% by mid-2026, which leaves considerable headroom.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory