GDPval
Also known as GDPval-AA, GDPval-AA v2
GDPval measures whether a model can produce the actual deliverables of skilled professional work, the documents, slides, spreadsheets and diagrams that people are paid for, well enough to stand against an industry expert's version of the same job. OpenAI built it because exam-style benchmarks stopped telling anyone whether AI could do work, and it is the closest thing the field has to an economic yardstick.
| What it measures | Whether a model can produce the actual deliverables of skilled professional work (documents, slides, spreadsheets, diagrams) well enough to stand against an industry expert's version. |
|---|---|
| Built by | OpenAI, with an agentic re-run by Artificial Analysis as GDPval-AA, 2025 |
| Format | 1,320 tasks spanning 44 occupations across the 9 largest sectors of US GDP, written by professionals averaging 14 years of experience; a 220-task gold subset is open source and is what the Artificial Analysis v2 run uses, with shell access and web browsing in an agentic loop |
| Scoring metric | Blind pairwise comparison of two anonymised outputs on the same task, aggregated into an Elo rating; the v2 scale anchors human expert deliverables at 1000 |
| Status | Active |
| Representative top score | 1,764 Elo (GDPval-AA v2) · Claude Fable 5.1 (adaptive reasoning, max effort) · read 2026-09 |
| Official leaderboard | artificialanalysis.ai/evaluations/gdpval-aa |
How GDPval works
GDPval covers the majority of US Bureau of Labor Statistics work activities for 44 occupations across the nine sectors contributing most to US GDP. The tasks were written by industry professionals averaging 14 years of experience, and the full set holds 1,320 of them; a gold subset of 220 is open source. Grading is blinded pairwise comparison: an expert in the relevant occupation sees the request, the reference files, and two or more unlabelled deliverables, and ranks them. The reported figure is how often a model output beats or matches the human expert version. An experimental automated grader reached 66% agreement with human experts, against 71% agreement between the humans themselves.
History and current status
OpenAI released GDPval in October 2025 (arXiv 2510.04374). Artificial Analysis then built GDPval-AA, an independent re-run of the 220-task gold subset that gives models shell access and web browsing in an agentic loop rather than asking for a single-shot deliverable, and scores the blind comparisons as an Elo rating anchored so that human expert work sits at 1000. GDPval-AA v2 is now a weighted component of the Artificial Analysis Intelligence Index, which is how the benchmark reaches most people who have never read the paper.
What the score does not tell you
Pairwise preference measures which deliverable a grader prefers, not which one is correct, and a polished wrong answer can beat a plain right one. The automated grader agrees with human experts less often than humans agree with each other, so any score produced at scale carries that gap. Task selection follows BLS work activities, which describe occupations rather than the specific jobs any given employer needs done. On the Artificial Analysis re-run the judge is itself a language model, which imports the well-documented preference of such judges for length and formatting. Top models now sit far above the 1000-Elo human anchor, and that should be read as preferred output, not verified correctness.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
GDPval is not one of the 15 benchmarks scored in the benchmark trust ranking, which rates the most-quoted tests on those same four axes and sorts them most trustworthy first.
What GDPval scores actually mean
The two scales are not interchangeable. On OpenAI's own protocol, the best model in the paper, Claude Opus 4.1, beat or matched human expert deliverables on 47.6% of tasks, with GPT-5 at 39.0%, o3 at 35.2% and GPT-4o at 12.5%. Read that as roughly half, not as superhuman. On the Artificial Analysis GDPval-AA v2 board, where human expert work is pinned at 1000 Elo, Claude Fable 5.1 at max effort read 1,764 and Claude Opus 5 at max effort read 1,735 when checked in September 2026. Those numbers are drifting, and not always upward: the same Opus 5 configuration read 1,861 in July 2026. An Elo far above 1000 does not mean the model is 76% better than a professional; it means blind graders picked its deliverable that much more often under an agentic harness the original paper did not use.
Who reports GDPval, and how to read it
OpenAI reports it in research posts rather than in launch marketing, and rival labs largely ignore it, which is unsurprising for a benchmark one lab designed. Its real distribution is through Artificial Analysis, both on the GDPval-AA leaderboard and as a 10% weight inside Intelligence Index v4.3. That means a great many people quote GDPval indirectly, through a composite, without knowing the component is there or that it is an agentic re-run rather than the original protocol. When a score is cited, ask which one it is: the OpenAI win rate against experts, or the Artificial Analysis Elo.
When to weight GDPval in a model choice
GDPval is the benchmark to reach for when the question is economic rather than technical, because it is the only entry here whose unit is professional output. The paper carries the figures that make it usable for costing: experts averaged 404 minutes per task, at an average task cost of $361 computed from median occupational wages. Set a model run against those two numbers and you have a defensible first estimate of what automating a category of work is worth. Do not use it to rank frontier models against each other, because the preference scale and the drifting Elo make small gaps meaningless, and always state which protocol a quoted score came from.
Benchmarks to read alongside this one
Artificial Analysis Intelligence Index
A composite index of overall model intelligence aggregating performance across reasoning, coding, knowledge, science and agentic tasks.
SWE-Lancer
Whether frontier models can complete real paid freelance software jobs, both coding and technical-management tasks, well enough to earn the payouts.
FinanceBench
Open-book financial question answering over real public-company filings, with the supporting evidence required.
GDPval: frequently asked questions
- What is GDPval?
- GDPval is a 2025 OpenAI benchmark that tests whether AI models can produce real professional deliverables across 44 occupations in the nine sectors contributing most to US GDP. It holds 1,320 tasks written by professionals averaging 14 years of experience, with a 220-task gold subset released open source.
- Have AI models beaten human experts on GDPval?
- Not on the original protocol. The best model in the paper, Claude Opus 4.1, beat or matched expert deliverables on 47.6% of tasks. On the separate Artificial Analysis GDPval-AA agentic re-run, top models score well above the 1000-Elo human anchor, but that scale measures blind grader preference, not verified correctness.
- What does a GDPval task cost a human to do?
- The paper reports that industry experts took an average of 404 minutes per task, roughly 6.7 hours, at an average cost of $361 computed from median occupational hourly wages. Those two figures are what make GDPval usable for costing automation.
Sources
- Patwardhan, Dias, Proehl, Kim, Wang, Watkins et al. (2025). GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks. arXiv preprint 2510.04374. arxiv.org/abs/2510.04374
- Artificial Analysis. GDPval-AA evaluation and leaderboard (Elo, human expert anchored at 1000). artificialanalysis.ai/evaluations/gdpval-aa
- Artificial Analysis (2026). Announcing the Artificial Analysis Intelligence Index v4.3, which weights GDPval-AA v2 at 10%. artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.
← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →