Capital & Compute
Benchmark· Reasoning & abstraction· Checked 2026-07-26

ARC-AGI-3

Also known as ARC-AGI v3, ARC-AGI 3

ARC-AGI-3 is the first fully interactive ARC benchmark. Instead of a static puzzle, an agent is dropped into an unfamiliar game environment with no instructions, no stated goal and no rules, and has to work out what to do by acting. It currently shows the widest human-model gap of any benchmark in this directory: humans scored 100% at its launch while frontier AI managed 0.51%.

Key facts about the ARC-AGI-3 benchmark
What it measuresWhether an agent dropped into an unfamiliar interactive environment with no instructions, stated goal or rules can work out what to do by acting, build a usable world model, and keep learning across levels.
Built byARC Prize Foundation, 2026
FormatHundreds of handcrafted interactive game environments spanning thousands of levels, played through an SDK, a REST API, or in the browser
Scoring metricGames beaten at or above human-level action efficiency, measuring skill-acquisition efficiency rather than one-shot accuracy
StatusActive
Representative top score30.2% (public demo) · Claude Opus 5 (high effort) · read 2026-07
Official leaderboardarcprize.org/leaderboard

How ARC-AGI-3 works

The benchmark comprises hundreds of handcrafted interactive game environments spanning thousands of levels, playable through an SDK, a REST API or a browser. The measured quantity is not one-shot accuracy but skill-acquisition efficiency: how many games the agent beats at or above human-level action efficiency. An agent that eventually stumbles into a solution after vastly more actions than a person does not score well.

History and current status

ARC Prize launched it on 25 March 2026, after ARC-AGI-2 fell far faster than expected. At launch the human-AI gap was almost total. Progress since has been rapid in relative terms and negligible in absolute terms: ARC Prize verified Claude Opus 5 at 30.16% on the 25-environment public demo in July 2026, roughly four times what GPT-5.6 Sol reached at maximum effort. ARC Prize 2026 carries over 2 million dollars in prizes.

What the score does not tell you

Interactive benchmarks are harder to standardise than static ones: results depend on the action budget, the harness and the effort setting, and the published comparisons are not always effort-matched. The Opus 5 figure, for instance, is a high-effort run and is not directly comparable to a different model at a different setting. The public demo is also only a slice of the full environment set.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

Who reports ARC-AGI-3, and how to read it

ARC Prize verifies and publishes the numbers, which makes this a rare independently operated frontier benchmark. Because the absolute scores are low and the setup matters, prefer ARC-Prize-verified figures with the environment split and effort level stated, and be sceptical of any round number quoted without those.

2026
First released
ARC Prize Foundation
Active
Status today
As of July 27, 2026
30.2% (public demo)
Representative top score
Read 2026-07

Benchmarks to read alongside this one

ARC-AGI-3: frequently asked questions

What is ARC-AGI-3?
ARC-AGI-3 is the first fully interactive ARC benchmark, launched on 25 March 2026. An agent enters an unfamiliar game environment with no instructions, goal or rules and must learn what to do by acting. It measures skill-acquisition efficiency against human action efficiency, not one-shot accuracy.
What is the hardest AI benchmark in 2026?
By the size of the human-model gap, ARC-AGI-3. Humans scored 100% at launch against 0.51% for frontier AI, and the ARC-Prize-verified top result was about 30% in July 2026. Humanity’s Last Exam is the hardest of the knowledge-style benchmarks, at roughly 53%.
Why is ARC-AGI-3 so much harder than ARC-AGI-2?
ARC-AGI-2 gives you a puzzle with visible examples. ARC-AGI-3 gives you an environment and tells you nothing: you have to discover the goal and the rules through interaction, then generalise across levels. That is exploration and world-model building, not pattern inference.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory