Capital & Compute
Benchmark· Reasoning & abstraction· Checked 2026-07-01

GPQA Diamond

Also known as GPQA-Diamond

GPQA Diamond is a set of 198 graduate-level questions in biology, physics and chemistry, written by domain experts and deliberately designed so that a skilled non-expert with a search engine still cannot answer them quickly. That "Google-proof" property is what made it the standard test of reasoning over retrieval. By 2026 frontier models exceed the human-expert baseline and it is effectively saturated.

Key facts about the GPQA Diamond benchmark
What it measuresGraduate and PhD-level multiple-choice scientific reasoning in biology, physics and chemistry, on questions designed to be unanswerable by quick web search.
Built byRein et al. (NYU, Cohere, Anthropic), 2023
Format198 expert-written four-option questions (the hardest, highest-agreement subset of the 448-question GPQA set)
Scoring metricMultiple-choice accuracy (random baseline 25%, PhD-expert baseline about 70%)
StatusSaturated
Representative top score~94% · Gemini 3.1 Pro Preview · read 2026-02
Official leaderboardepoch.ai/benchmarks/gpqa-diamond

How GPQA Diamond works

Diamond is the hardest subset of the 448-question GPQA set: the items where the writing experts agreed on the answer and validating experts in other fields failed to get it right even with web access. Each question is four-option multiple choice, so the random baseline is 25%, and the PhD-expert baseline is around 70%. Scoring is plain accuracy.

History and current status

Introduced in a 2023 paper by Rein and collaborators at NYU with Cohere and Anthropic, it sat well below the expert baseline when reasoning models arrived and then became the benchmark that demonstrated their advantage. Scores passed the human-expert level and by early 2026 the top of the range is in the low-to-mid 90s, which is why it now sits in this directory as saturated rather than active.

What the score does not tell you

The most practical problem is size. With only 198 items, a handful of questions moves the headline number by a full percentage point, so small reported differences between models are statistically meaningless. Beyond that, the questions have circulated for years, and above the human-expert baseline the remaining margin is partly a measure of the answer key rather than of understanding.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, why benchmark saturation makes a high score meaningless.

Who reports GPQA Diamond, and how to read it

GPQA Diamond is still a fixture of frontier launch tables. Epoch AI runs an independent tracker, which is the figure to prefer over a self-reported one. Where secondary coverage cites a very high number that Epoch has not corroborated, treat it as unverified: this directory holds its recorded top score at roughly 94% for that reason.

2023
First released
Rein et al. (NYU, Cohere, Anthropic)
Saturated
Status today
As of July 27, 2026
~94%
Representative top score
Read 2026-02

Benchmarks to read alongside this one

GPQA Diamond: frequently asked questions

What is GPQA Diamond?
GPQA Diamond is a 198-question multiple-choice benchmark of graduate and PhD-level biology, physics and chemistry problems written by domain experts to be Google-proof, meaning a non-expert with web access cannot answer them quickly. It tests reasoning rather than retrieval.
What does Google-proof mean?
It means the question was validated by having skilled people in other fields attempt it with full web access. Only questions they still got wrong were kept, so a correct answer cannot come from a quick search and has to come from actual domain reasoning.
Is GPQA Diamond saturated?
Yes, effectively. Top models now sit in the low-to-mid 90s against a PhD-expert baseline of about 70%. Combined with only 198 items, the differences between frontier models fall within noise.
How many questions are in GPQA Diamond?
198. It is the hardest, highest-expert-agreement subset of the full 448-question GPQA set. That small size is its main statistical weakness: a few items swing the reported score.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory