Capital & Compute
Benchmark· Reasoning & abstraction· Checked 2026-07-26

ARC-AGI-2

Also known as ARC-AGI v2

ARC-AGI-2 is a test of fluid intelligence: grid puzzles where you must infer a novel rule from a few examples and apply it, rather than recall anything. It was designed to be easy for humans and very hard for AI, and it held that property for about a year before frontier models cleared it. Its collapse in 2026 is one of the cleanest saturation stories in this directory.

Key facts about the ARC-AGI-2 benchmark
What it measuresThe same fluid-intelligence test as ARC-AGI-1, but with harder, contamination-resistant tasks that stay easy for humans yet very hard for AI.
Built byARC Prize Foundation (Chollet et al.), 2025
Format1,240 grid tasks (1,000 training, 120 public, 120 semi-private, 120 private eval), each solvable pass@2 by at least two humans
Scoring metricpass@2 exact-grid-match accuracy, reported with a cost-per-task efficiency metric
StatusSaturated
Representative top score92.5% (semi-private) · GPT-5.6 Sol (max effort) · read 2026-07
Official leaderboardarcprize.org/leaderboard

How ARC-AGI-2 works

The benchmark contains 1,240 grid tasks split into training, public evaluation, semi-private evaluation and private evaluation sets. Every task was validated as solvable within two attempts by at least two humans. Scoring is pass@2 exact grid match: the output grid is either exactly right or wrong. Crucially, results are reported alongside a cost-per-task figure, because a model can brute-force accuracy by spending far more compute.

History and current status

The ARC Prize Foundation released it in 2025 as the successor to ARC-AGI-1, which frontier systems had reached 97.5% on. As recently as December 2025 the best verified score was around 54%. By July 2026 ARC-Prize-verified runs put GPT-5.6 Sol at 92.5% on the semi-private set and Claude Opus 5 at 90.4%, comfortably past the 85% target. Attention has moved to ARC-AGI-3.

What the score does not tell you

The benchmark design is sound: the semi-private and private splits are unpublished, so contamination stays genuinely low, and only ARC-Prize-verified numbers should be trusted rather than self-reported ones. The criticism is about what saturation means here. Clearing a fluid-reasoning test that took a year to fall does not settle whether the underlying capability is general, and cost per task is now the more informative column than accuracy.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, why benchmark saturation makes a high score meaningless.

Who reports ARC-AGI-2, and how to read it

ARC Prize runs verification itself and publishes a leaderboard, so this is one of the few benchmarks where an independent operator controls the number. Labs cite ARC-AGI figures at launch; check whether the figure is ARC-Prize-verified and which split it refers to, since public, semi-private and private results differ.

2025
First released
ARC Prize Foundation (Chollet et al.)
Saturated
Status today
As of July 27, 2026
92.5% (semi-private)
Representative top score
Read 2026-07

Benchmarks to read alongside this one

ARC-AGI-2: frequently asked questions

What is ARC-AGI-2?
ARC-AGI-2 is a 2025 benchmark of 1,240 grid-puzzle tasks that test fluid reasoning: inferring a novel rule from a few examples and applying it. Every task was verified solvable within two attempts by at least two humans, and scoring is exact grid match at pass@2.
Has ARC-AGI-2 been solved?
Effectively yes. ARC-Prize-verified runs reached 92.5% on the semi-private set by July 2026, past the 85% target, up from about 54% in December 2025. It no longer separates frontier models, which is why ARC-AGI-3 exists.
Why is cost per task reported with ARC-AGI scores?
Because accuracy on these puzzles can be bought with compute. A model can improve by searching far longer, so an accuracy figure without a cost figure hides how the score was achieved. ARC Prize reports both for that reason.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory