GPQA Diamond
Also known as GPQA-Diamond
GPQA Diamond is a set of 198 graduate-level questions in biology, physics and chemistry, written by domain experts and deliberately designed so that a skilled non-expert with a search engine still cannot answer them quickly. That "Google-proof" property is what made it the standard test of reasoning over retrieval. By 2026 frontier models exceed the human-expert baseline and it is effectively saturated.
| What it measures | Graduate and PhD-level multiple-choice scientific reasoning in biology, physics and chemistry, on questions designed to be unanswerable by quick web search. |
|---|---|
| Built by | Rein et al. (NYU, Cohere, Anthropic), 2023 |
| Format | 198 expert-written four-option questions (the hardest, highest-agreement subset of the 448-question GPQA set) |
| Scoring metric | Multiple-choice accuracy (random baseline 25%, PhD-expert baseline about 70%) |
| Status | Saturated |
| Representative top score | ~94% · Gemini 3.1 Pro Preview · read 2026-02 |
| Official leaderboard | epoch.ai/benchmarks/gpqa-diamond |
How GPQA Diamond works
Diamond is the hardest subset of the 448-question GPQA set: the items where the writing experts agreed on the answer and validating experts in other fields failed to get it right even with web access. Each question is four-option multiple choice, so the random baseline is 25%, and the PhD-expert baseline is around 70%. Scoring is plain accuracy.
History and current status
Introduced in a 2023 paper by Rein and collaborators at NYU with Cohere and Anthropic, it sat well below the expert baseline when reasoning models arrived and then became the benchmark that demonstrated their advantage. Scores passed the human-expert level and by early 2026 the top of the range is in the low-to-mid 90s, which is why it now sits in this directory as saturated rather than active.
What the score does not tell you
The most practical problem is size. With only 198 items, a handful of questions moves the headline number by a full percentage point, so small reported differences between models are statistically meaningless. Beyond that, the questions have circulated for years, and above the human-expert baseline the remaining margin is partly a measure of the answer key rather than of understanding.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, why benchmark saturation makes a high score meaningless.
Who reports GPQA Diamond, and how to read it
GPQA Diamond is still a fixture of frontier launch tables. Epoch AI runs an independent tracker, which is the figure to prefer over a self-reported one. Where secondary coverage cites a very high number that Epoch has not corroborated, treat it as unverified: this directory holds its recorded top score at roughly 94% for that reason.
Benchmarks to read alongside this one
Humanity's Last Exam
Frontier, closed-ended expert knowledge and reasoning across more than 100 academic disciplines at the limit of human expertise.
SuperGPQA
Graduate-level knowledge and reasoning across 285 disciplines, including the applied and service fields that mainstream benchmarks ignore.
MMLU-Pro
Harder multi-task reasoning and knowledge designed to de-saturate MMLU and reward deliberate reasoning over recall.
GPQA Diamond: frequently asked questions
- What is GPQA Diamond?
- GPQA Diamond is a 198-question multiple-choice benchmark of graduate and PhD-level biology, physics and chemistry problems written by domain experts to be Google-proof, meaning a non-expert with web access cannot answer them quickly. It tests reasoning rather than retrieval.
- What does Google-proof mean?
- It means the question was validated by having skilled people in other fields attempt it with full web access. Only questions they still got wrong were kept, so a correct answer cannot come from a quick search and has to come from actual domain reasoning.
- Is GPQA Diamond saturated?
- Yes, effectively. Top models now sit in the low-to-mid 90s against a PhD-expert baseline of about 70%. Combined with only 198 items, the differences between frontier models fall within noise.
- How many questions are in GPQA Diamond?
- 198. It is the hardest, highest-expert-agreement subset of the full 448-question GPQA set. That small size is its main statistical weakness: a few items swing the reported score.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.