ARC-AGI-3
Also known as ARC-AGI v3, ARC-AGI 3
ARC-AGI-3 is the first fully interactive ARC benchmark. Instead of a static puzzle, an agent is dropped into an unfamiliar game environment with no instructions, no stated goal and no rules, and has to work out what to do by acting. It currently shows the widest human-model gap of any benchmark in this directory: humans scored 100% at its launch while frontier AI managed 0.51%.
| What it measures | Whether an agent dropped into an unfamiliar interactive environment with no instructions, stated goal or rules can work out what to do by acting, build a usable world model, and keep learning across levels. |
|---|---|
| Built by | ARC Prize Foundation, 2026 |
| Format | Hundreds of handcrafted interactive game environments spanning thousands of levels, played through an SDK, a REST API, or in the browser |
| Scoring metric | Games beaten at or above human-level action efficiency, measuring skill-acquisition efficiency rather than one-shot accuracy |
| Status | Active |
| Representative top score | 30.2% (public demo) · Claude Opus 5 (high effort) · read 2026-07 |
| Official leaderboard | arcprize.org/leaderboard |
How ARC-AGI-3 works
The benchmark comprises hundreds of handcrafted interactive game environments spanning thousands of levels, playable through an SDK, a REST API or a browser. The measured quantity is not one-shot accuracy but skill-acquisition efficiency: how many games the agent beats at or above human-level action efficiency. An agent that eventually stumbles into a solution after vastly more actions than a person does not score well.
History and current status
ARC Prize launched it on 25 March 2026, after ARC-AGI-2 fell far faster than expected. At launch the human-AI gap was almost total. Progress since has been rapid in relative terms and negligible in absolute terms: ARC Prize verified Claude Opus 5 at 30.16% on the 25-environment public demo in July 2026, roughly four times what GPT-5.6 Sol reached at maximum effort. ARC Prize 2026 carries over 2 million dollars in prizes.
What the score does not tell you
Interactive benchmarks are harder to standardise than static ones: results depend on the action budget, the harness and the effort setting, and the published comparisons are not always effort-matched. The Opus 5 figure, for instance, is a high-effort run and is not directly comparable to a different model at a different setting. The public demo is also only a slice of the full environment set.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Who reports ARC-AGI-3, and how to read it
ARC Prize verifies and publishes the numbers, which makes this a rare independently operated frontier benchmark. Because the absolute scores are low and the setup matters, prefer ARC-Prize-verified figures with the environment split and effort level stated, and be sceptical of any round number quoted without those.
Benchmarks to read alongside this one
ARC-AGI-2
The same fluid-intelligence test as ARC-AGI-1, but with harder, contamination-resistant tasks that stay easy for humans yet very hard for AI.
OSWorld 2.0
Whether a computer-use agent can finish long-horizon professional work on a real desktop, coordinating several applications and leaving the machine in the correct final state.
METR Time Horizon
Model capability expressed in human time: the length of task, measured by how long humans take, that a model completes with 50% success.
ARC-AGI-3: frequently asked questions
- What is ARC-AGI-3?
- ARC-AGI-3 is the first fully interactive ARC benchmark, launched on 25 March 2026. An agent enters an unfamiliar game environment with no instructions, goal or rules and must learn what to do by acting. It measures skill-acquisition efficiency against human action efficiency, not one-shot accuracy.
- What is the hardest AI benchmark in 2026?
- By the size of the human-model gap, ARC-AGI-3. Humans scored 100% at launch against 0.51% for frontier AI, and the ARC-Prize-verified top result was about 30% in July 2026. Humanity’s Last Exam is the hardest of the knowledge-style benchmarks, at roughly 53%.
- Why is ARC-AGI-3 so much harder than ARC-AGI-2?
- ARC-AGI-2 gives you a puzzle with visible examples. ARC-AGI-3 gives you an environment and tells you nothing: you have to discover the goal and the rules through interaction, then generalise across levels. That is exploration and world-model building, not pattern inference.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.