Capital & Compute

ARC-AGI-1

Benchmark· Reasoning & abstraction· Checked 2026-09-12

Also known as ARC-AGI, Abstraction and Reasoning Corpus

ARC-AGI-1 is the grid-puzzle benchmark Francois Chollet built in 2019 to test skill acquisition rather than knowledge. Each task shows about three example input-output grids and asks for the output of a fourth, so nothing can be looked up. It held out against AI for five years, which is far longer than anything else in this directory, and its fall in late 2024 is the single clearest marker of what changed at the frontier.

Key facts about the ARC-AGI-1 benchmark
What it measuresWhether a system can infer the abstract rule of a novel visual grid puzzle from a few examples and apply it to a new input.
Built byFrancois Chollet (ARC Prize Foundation), 2019
Format800 public tasks (400 training, 400 public eval) plus a 100-task semi-private and a 100-task private eval set; each task gives about three input-output grid examples and a test input
Scoring metricpass@2 exact-grid-match accuracy
StatusSaturated
Representative top score97.5% (public eval) · Claude Opus 5 and GPT-5.6 Sol · read 2026-07
Official leaderboardarcprize.org/leaderboard

How ARC-AGI-1 works

A task is a handful of coloured grids. The model sees roughly three input-output pairs that demonstrate some abstract transformation, then a test input, and must produce the exact output grid. There is no partial credit: the grid matches or it does not, and the standard scoring is pass@2, meaning two attempts are allowed. The corpus holds 800 public tasks, split 400 for training and 400 for public evaluation, plus a 100-task semi-private set and a 100-task private set that exist so a score cannot be obtained by memorising published answers. Every task is designed to be solvable by a human with no special training and to require a rule no model could have seen before.

History and current status

Chollet introduced the Abstraction and Reasoning Corpus in 2019 alongside his paper On the Measure of Intelligence, which argued that intelligence should be measured as skill acquisition efficiency rather than as accumulated skill. It resisted every approach through the deep learning boom. In December 2024 OpenAI's o3-preview cleared it, which ARC Prize reported as 75.7% at low compute and 87.5% at high compute, passing the 85% target that had stood since launch. By 2026 the public eval set is effectively finished, with ARC-Prize-verified runs reported in the mid-to-high nineties. ARC-AGI-2 followed in 2025 and ARC-AGI-3, an interactive version, launched in March 2026.

What the score does not tell you

The pass@2 exact-match rule is unforgiving in a way that flatters nothing and punishes near-misses equally with nonsense, which makes the score harder to interpret than an accuracy figure. The compute caveat is the bigger issue: the 2024 breakthrough arrived at a cost per task high enough that ARC Prize reported low-compute and high-compute figures separately, so a headline percentage without a cost figure hides most of what happened. Public eval answers have been available for years and are certainly in training corpora, which is exactly why the semi-private and private sets exist, and why any score quoted against the public 400 should be treated as contaminated by default.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, why benchmark saturation makes a high score meaningless.

ARC-AGI-1 is not one of the 15 benchmarks scored in the benchmark trust ranking, which rates the most-quoted tests on those same four axes and sorts them most trustworthy first.

What ARC-AGI-1 scores actually mean

Read ARC-AGI-1 as history now rather than as a live signal, because the top of the board sits in the mid-to-high nineties and no longer separates frontier models. The informative band was always lower down. Through 2020 to 2023 the best program-synthesis approaches sat in the teens and twenties while humans solved the same tasks comfortably, and that gap was the whole argument: a system could pass professional exams and still fail puzzles a child handles. When o3-preview jumped to 75.7% and then 87.5% in one release, the interesting number was not the percentage but the compute it took to get there. Today a score below about 90% on the public set signals a smaller or non-reasoning model, and anything above that tells you almost nothing, which is why ARC-AGI-2 and ARC-AGI-3 exist.

Who reports ARC-AGI-1, and how to read it

ARC Prize verifies and publishes runs itself, which is unusual and valuable: most benchmarks here rely on self-reported lab numbers, whereas an ARC-verified score has been reproduced by the foundation on a held-out set. Labs quote ARC-AGI in launch material when they clear a threshold, and almost always against the public eval rather than the private set. The gap between a self-reported public-eval number and an ARC-verified semi-private number is the thing to look for, and it has been large enough in past cases to reverse a claim.

When to weight ARC-AGI-1 in a model choice

Do not use ARC-AGI-1 to choose between current frontier models; it is saturated and will not separate them. Use it for two other things. First, as the reference point for how fast a benchmark can go from unbeatable to finished, which is five years here against roughly two for most of the coding sets. Second, as a check on reasoning claims in smaller or open-weight models, where the public eval still discriminates. If you want a live abstraction signal, move to ARC-AGI-2, or to ARC-AGI-3 if you care about interactive exploration rather than single-shot rule inference.

2019
First released
Francois Chollet (ARC Prize Foundation)
Saturated
Status today
As of September 12, 2026
97.5% (public eval)
Representative top score
Read 2026-07

Benchmarks to read alongside this one

ARC-AGI-1: frequently asked questions

What is ARC-AGI-1?
ARC-AGI-1 is a 2019 benchmark from Francois Chollet that tests whether a system can infer the abstract rule behind a few example grid transformations and apply it to a new input. It holds 800 public tasks plus 100-task semi-private and private eval sets, scored pass@2 on exact grid match.
Has ARC-AGI-1 been solved?
Effectively yes. It resisted AI from 2019 until December 2024, when OpenAI o3-preview cleared the 85% target, reported by ARC Prize as 75.7% at low compute and 87.5% at high compute. Frontier scores now sit in the mid-to-high nineties, which is why ARC-AGI-2 and ARC-AGI-3 were built.
What is the difference between ARC-AGI-1, 2 and 3?
ARC-AGI-1 is single-shot grid rule inference and is saturated. ARC-AGI-2 raised the difficulty on the same format. ARC-AGI-3, launched March 2026, is interactive: the system must explore an environment rather than read a fixed set of examples, and frontier models score close to zero on it.

Sources

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →