GAIA
Also known as General AI Assistants benchmark
GAIA tests whether an AI assistant can answer questions that are conceptually simple for a person but require real work: browsing the web, handling several file types, and chaining multiple steps. Each question has one unambiguous answer, so grading is exact-match. It became a reference point because the answers to its test set are private.
| What it measures | Whether an AI assistant can answer real-world questions that require multi-step reasoning, multiple modalities, web browsing and general tool use. |
|---|---|
| Built by | Meta AI and Hugging Face (Mialon, Fourrier et al.), 2023 |
| Format | 466 real-world questions across 3 difficulty levels (165 public validation, about 300 held-out test), each needing tools or browsing and a single answer |
| Scoring metric | Exact-match accuracy against an unambiguous answer |
| Status | Active |
| Representative top score | ~75% · HAL agent (Claude Sonnet 4.5) · read 2026-06 |
| Official leaderboard | hal.cs.princeton.edu/gaia |
How GAIA works
The benchmark contains 466 real-world questions across three difficulty levels, split into roughly 165 public validation questions and about 300 held-out test questions. Every question requires tool use or browsing and resolves to a single unambiguous answer, which is why grading can be exact-match rather than judged. Test-set submissions are graded by the operator, so the answers never become public.
History and current status
Researchers at Meta AI and Hugging Face introduced GAIA in 2023, and it became the standard general-assistant benchmark through the first wave of agent products. Princeton HAL leaderboard later reframed it around agent reliability and cost rather than raw accuracy alone, which matters because a high GAIA score achieved with an enormous number of tool calls is a different product from the same score achieved cheaply.
What the score does not tell you
The private test set limits contamination but not drift: questions depend on the live web, and web pages change, so a question that was answerable in 2023 may not be answerable the same way now. Exact-match grading on a single answer also penalises a correct answer expressed differently, and rewards agents tuned to the expected output format.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Who reports GAIA, and how to read it
GAIA appears in agent-product announcements and on the Princeton HAL leaderboard, which is the better source because it reports cost alongside accuracy. When reading a GAIA figure, check whether it is validation or test, since the public validation set can be optimised against and the held-out test set cannot.
Benchmarks to read alongside this one
Mind2Web 2
Whether an agentic search or deep-research system can browse the live web and return a correct, citation-backed answer to a long-horizon question.
tau-bench
Whether a tool-using agent can reliably complete customer-service tasks over multi-turn conversations with a simulated user while obeying domain policies.
BrowseComp
Whether a browsing agent can persistently navigate the open web to locate a single hard-to-find, entangled fact.
GAIA: frequently asked questions
- What is the GAIA benchmark?
- GAIA is a 2023 benchmark of 466 real-world questions that require multi-step reasoning, several modalities, web browsing and tool use. Roughly 165 questions form a public validation set and about 300 are held out as a private test set, each with one unambiguous answer.
- Why is GAIA hard for AI but easy for humans?
- Because the difficulty is in execution, not comprehension. A person can follow a chain of steps across a few websites and file formats without difficulty; an agent has to browse reliably, parse different formats, and keep track of intermediate results without losing the thread.
- Is GAIA contaminated?
- Less than most benchmarks, because the test-set answers are private and submissions are graded by the operator. The residual problem is drift rather than leakage: the questions depend on the live web, which changes underneath them.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.