Capital & Compute

GAIA

Benchmark· Agents, tool use & computer use· Checked 2026-06-29

Also known as General AI Assistants benchmark

GAIA tests whether an AI assistant can answer questions that are conceptually simple for a person but require real work: browsing the web, handling several file types, and chaining multiple steps. Each question has one unambiguous answer, so grading is exact-match. It became a reference point because the answers to its test set are private.

Key facts about the GAIA benchmark
What it measuresWhether an AI assistant can answer real-world questions that require multi-step reasoning, multiple modalities, web browsing and general tool use.
Built byMeta AI and Hugging Face (Mialon, Fourrier et al.), 2023
Format466 real-world questions across 3 difficulty levels (165 public validation, about 300 held-out test), each needing tools or browsing and a single answer
Scoring metricExact-match accuracy against an unambiguous answer
StatusActive
Representative top score~75% · HAL agent (Claude Sonnet 4.5) · read 2026-06
Official leaderboardhal.cs.princeton.edu/gaia

How GAIA works

The benchmark contains 466 real-world questions across three difficulty levels, split into roughly 165 public validation questions and about 300 held-out test questions. Every question requires tool use or browsing and resolves to a single unambiguous answer, which is why grading can be exact-match rather than judged. Test-set submissions are graded by the operator, so the answers never become public.

History and current status

Researchers at Meta AI and Hugging Face introduced GAIA in 2023, and it became the standard general-assistant benchmark through the first wave of agent products. Princeton HAL leaderboard later reframed it around agent reliability and cost rather than raw accuracy alone, which matters because a high GAIA score achieved with an enormous number of tool calls is a different product from the same score achieved cheaply.

What the score does not tell you

The private test set limits contamination but not drift: questions depend on the live web, and web pages change, so a question that was answerable in 2023 may not be answerable the same way now. Exact-match grading on a single answer also penalises a correct answer expressed differently, and rewards agents tuned to the expected output format.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

What GAIA scores actually mean

GAIA reports across three difficulty levels and a single averaged number hides most of what matters. Level 1 questions need a few tool calls; Level 3 questions need long chains where one wrong step ends the attempt. A system at 75% overall is typically near-ceiling on Level 1 and well under half on Level 3, so the average flatters. The held-out test answers are private and graded on submission, which keeps contamination low and makes GAIA figures more trustworthy than most agent benchmarks. Scores also depend heavily on the scaffolding wrapped around the model, so a GAIA number describes an agent system rather than a model, and swapping the harness moves it substantially.

Who reports GAIA, and how to read it

GAIA appears in agent-product announcements and on the Princeton HAL leaderboard, which is the better source because it reports cost alongside accuracy. When reading a GAIA figure, check whether it is validation or test, since the public validation set can be optimised against and the held-out test set cannot. Any figure should also name the scaffolding, because GAIA scores describe an agent system rather than a bare model.

When to weight GAIA in a model choice

Read GAIA when evaluating an assistant that has to browse, use tools and return one verifiable answer, which is a narrower and more testable claim than general agentic ability. Insist on the per-level breakdown rather than the average, because Level 3 is where the difference between a demo and a usable system shows. The Princeton HAL board is the better source now, since it reports reliability and cost alongside accuracy, and cost is what determines whether an agent that eventually gets the answer is worth running. Treat any GAIA figure as a property of the full agent stack, not of the underlying model.

2023
First released
Meta AI and Hugging Face (Mialon, Fourrier et al.)
Active
Status today
As of September 4, 2026
~75%
Representative top score
Read 2026-06

Benchmarks to read alongside this one

GAIA: frequently asked questions

What is the GAIA benchmark?
GAIA is a 2023 benchmark of 466 real-world questions that require multi-step reasoning, several modalities, web browsing and tool use. Roughly 165 questions form a public validation set and about 300 are held out as a private test set, each with one unambiguous answer.
Why is GAIA hard for AI but easy for humans?
Because the difficulty is in execution, not comprehension. A person can follow a chain of steps across a few websites and file formats without difficulty; an agent has to browse reliably, parse different formats, and keep track of intermediate results without losing the thread.
Is GAIA contaminated?
Less than most benchmarks, because the test-set answers are private and submissions are graded by the operator. The residual problem is drift rather than leakage: the questions depend on the live web, which changes underneath them.

Sources

  • Mialon, Fourrier, Swift, Wolf et al. (2023). GAIA: a benchmark for General AI Assistants. arXiv preprint 2311.12983. arxiv.org/abs/2311.12983
  • Princeton HAL. GAIA leaderboard reporting accuracy alongside agent cost and reliability. hal.cs.princeton.edu/gaia

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory