ARC-AGI-3
Also known as ARC-AGI v3, ARC-AGI 3
ARC-AGI-3 is the first fully interactive ARC benchmark. Instead of a static puzzle, an agent is dropped into an unfamiliar game environment with no instructions, no stated goal and no rules, and has to work out what to do by acting. It currently shows the widest human-model gap of any benchmark in this directory: humans scored 100% at its launch while frontier AI managed 0.51%.
| What it measures | Whether an agent dropped into an unfamiliar interactive environment with no instructions, stated goal or rules can work out what to do by acting, build a usable world model, and keep learning across levels. |
|---|---|
| Built by | ARC Prize Foundation, 2026 |
| Format | Hundreds of handcrafted interactive game environments spanning thousands of levels, played through an SDK, a REST API, or in the browser |
| Scoring metric | Games beaten at or above human-level action efficiency, measuring skill-acquisition efficiency rather than one-shot accuracy |
| Status | Active |
| Representative top score | 62.7% (semi-private, standard harness) · GPT-6 Astra (max effort) · read 2026-09 |
| Official leaderboard | arcprize.org/leaderboard |
How ARC-AGI-3 works
The benchmark comprises hundreds of handcrafted interactive game environments spanning thousands of levels, playable through an SDK, a REST API or a browser. The measured quantity is not one-shot accuracy but skill-acquisition efficiency: how many games the agent beats at or above human-level action efficiency. An agent that eventually stumbles into a solution after vastly more actions than a person does not score well.
History and current status
ARC Prize launched it on 25 March 2026, after ARC-AGI-2 fell far faster than expected. At launch the human-AI gap was almost total. Progress since has been rapid in relative terms and negligible in absolute terms: ARC Prize verified Claude Opus 5 at 30.16% on the 25-environment public demo in July 2026, against 13.33% for GPT-5.6 Sol at maximum effort on the same set. Watch which set a figure refers to, since Sol's widely quoted 7.78% is the semi-private number rather than the public one. From August 2026 the harness became the bigger variable: unverified third-party scaffolds began self-reporting public-set scores above 95%, three times the verified ceiling. ARC Prize 2026 carries over 2 million dollars in prizes.
What the score does not tell you
Interactive benchmarks are harder to standardise than static ones: results depend on the action budget, the harness and the effort setting, and the published comparisons are not always effort-matched. The Opus 5 figure, for instance, is a high-effort run and is not directly comparable to a different model at a different setting. The public demo is also only a slice of the full environment set. By August 2026 the harness problem had become the dominant one: two third-party harnesses self-reported public-set scores far above anything ARC Prize has verified, with Schema at 98.98% using Claude Opus 4.8 and Fable 5 on 15 July and Prime Agent at 95.5% using Claude Opus 5 on 5 August, against a verified ceiling of 30.2%. Neither has been checked by ARC Prize, so the public set now looks close to saturated by agent scaffolding while remaining unsolved by bare models. Treat any ARC-AGI-3 figure above 30.2% as a harness result until ARC verifies it.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Scored on those four axes, ARC-AGI-3 carries a concern score of 3 out of 8 and ranks 7 of 15, which puts it in the group that needs a caveat beside the number. See the reasoning behind that rating and how it compares with the rest of the field.
What ARC-AGI-3 scores actually mean
ARC-AGI-3 carries the widest human-model gap in this directory, and that gap is the number worth quoting. At its March 2026 launch humans scored 100% while frontier models managed 0.51%. Verified runs now put the leader at 30.2% on the 25-environment public demo, against 13.3% for the next system at max effort. Two cautions apply to every figure. The splits differ sharply: the widely circulated 7.78% for one model is a semi-private result, not the public-demo one, and the two are routinely quoted as if interchangeable. The runs are also not effort-matched, so a high-effort result compared against a max-effort result from a different lab is not a like-for-like comparison.
Who reports ARC-AGI-3, and how to read it
ARC Prize verifies and publishes the numbers, which makes this a rare independently operated frontier benchmark. Because the absolute scores are low and the setup matters, prefer ARC-Prize-verified figures with the environment split and effort level stated, and be sceptical of any round number quoted without those. Harness vendors also publish their own ARC-AGI-3 results, and as of August 2026 none of those has been verified by ARC Prize, so check whether a figure is verified or self-reported before comparing it to anything.
When to weight ARC-AGI-3 in a model choice
Watch ARC-AGI-3 as the leading indicator for interactive, skill-acquisition capability, the thing that most distinguishes an agent that can learn a new environment from one that pattern-matches a familiar one. It is not yet a model-selection input, because nothing scores highly enough for the ranking to bear weight on a purchasing decision. When citing it, always name the split and the effort setting; those two qualifiers account for more of the variance between quoted figures than the models do. For agentic capability you can act on today, use Terminal-Bench or OSWorld 2.0 instead.
Benchmarks to read alongside this one
ARC-AGI-2
The same fluid-intelligence test as ARC-AGI-1, but with harder, contamination-resistant tasks that stay easy for humans yet very hard for AI.
OSWorld 2.0
Whether a computer-use agent can finish long-horizon professional work on a real desktop, coordinating several applications and leaving the machine in the correct final state.
METR Time Horizon
Model capability expressed in human time: the length of task, measured by how long humans take, that a model completes with 50% success.
ARC-AGI-3: frequently asked questions
- What is ARC-AGI-3?
- ARC-AGI-3 is the first fully interactive ARC benchmark, launched on 25 March 2026. An agent enters an unfamiliar game environment with no instructions, goal or rules and must learn what to do by acting. It measures skill-acquisition efficiency against human action efficiency, not one-shot accuracy.
- What is the hardest AI benchmark in 2026?
- By the size of the human-model gap, ARC-AGI-3. Humans scored 100% at launch against 0.51% for frontier AI, and the ARC-Prize-verified top result was about 30% in July 2026. Humanity’s Last Exam is the hardest of the knowledge-style benchmarks, at roughly 53%.
- Why is ARC-AGI-3 so much harder than ARC-AGI-2?
- ARC-AGI-2 gives you a puzzle with visible examples. ARC-AGI-3 gives you an environment and tells you nothing: you have to discover the goal and the rules through interaction, then generalise across levels. That is exploration and world-model building, not pattern inference.
Sources
- ARC Prize Foundation (2026). ARC-AGI-3 interactive reasoning benchmark. arXiv preprint 2603.24621. arxiv.org/abs/2603.24621
- ARC Prize Foundation. Official verified leaderboard, listing splits and effort settings. arcprize.org/leaderboard
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.
← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →