Capital & Compute

Is GPT-6 Astra AGI? The 10 Tests That Decide It

OpenAI says the AGI era has started. Ten published tests, from ARC-AGI-3 to the market reaction, scored against the evidence with an honest tally.

· ai· benchmarks· agi· economics· By Capital & Compute

No. Not on any test anyone has actually published a result for.

OpenAI released GPT-6 Astra on September 3, 2026, and president Greg Brockman closed the briefing with “Welcome to the AGI era.” Three days later Nvidia’s Jensen Huang posted that AGI had arrived. Below are ten tests that could settle the question, each one scored against a bar somebody else set: the benchmark authors, a peer-reviewed taxonomy, OpenAI’s own charter, an independent evaluator, the bond market. Astra passes none of them outright. It comes close on two.

# Test Bar set by Verdict
1 ARC-AGI-3 saturation ARC Prize Foundation Partial
2 Does the score survive a harness change ARC Prize Foundation Fail
3 Placement on the Levels of AGI ladder Morris et al., Google DeepMind No reading
4 Autonomous task length METR No reading
5 Economically valuable work OpenAI’s own charter Partial
6 Do the independent aggregates agree Artificial Analysis, Epoch AI Fail
7 Does it know when it does not know Artificial Analysis Fail
8 Do the measurers endorse the claim ARC Prize, Chollet, Marcus, LeCun Fail
9 Have long real interest rates moved Chow, Halperin and Mazlish Fail
10 Has the equity market repriced The tape Fail

Two partials, two tests with no published reading at all, six failures, zero passes.

That tally is not a verdict on the model. Astra is the strongest system anyone has shipped, and on one measure it is genuinely startling. It is a verdict on the claim.

What would count as AGI, and who gets to say

The reason “is it AGI” produces such bad arguments is that the term has three incompatible definitions in active use, and the people making claims rarely say which one they mean.

OpenAI’s founding definition is economic: highly autonomous systems that outperform humans at most economically valuable work. That is a labor-market statement, testable in principle and nowhere near satisfied.

The second definition was financial, and it is dead. Microsoft and OpenAI’s contract at one point pegged AGI to a profit threshold reported at $100 billion, which meant the most consequential AGI declaration in the industry was a revenue milestone. That clause mutated from a capability test in 2018, to the profit number in December 2024, to a procedural sunset, and was removed entirely in April 2026 when Microsoft took a non-exclusive license through 2032. Nobody has replaced it with anything.

The third is the only one with a published rubric behind it. In Levels of AGI for Operationalizing Progress on the Path to AGI, Meredith Ringel Morris, Jascha Sohl-Dickstein, Shane Legg and colleagues at Google DeepMind (arXiv preprint, November 2023; ICML 2024 position paper) split the question into two axes that people constantly conflate. Performance is depth: how good, measured against percentiles of skilled adults. Generality is breadth: how many kinds of task. A chess engine is Superhuman on one axis and Narrow on the other, and calling it intelligent or not intelligent depends entirely on which axis you were looking at.

Their ladder runs Emerging (a match for an unskilled human), Competent (50th percentile of skilled adults), Expert (90th), Exceptional (99th), Superhuman (better than every human). Version 5 of the paper renamed the fourth rung from Virtuoso to Exceptional, which is worth knowing if you are comparing against older write-ups that use the old word.

The useful part is that Morris and colleagues placed the rival definitions on their own ladder. Aguera y Arcas and Norvig’s definition, they write, “would fall into the ‘Emerging AGI’ category of our ontology,” which is a bar frontier models cleared years ago. And OpenAI’s own threshold of labor replacement “better matches ‘Exceptional AGI’”: the 99th percentile of skilled adults, two full rungs above the level nobody has reached yet.

So when OpenAI’s president says the AGI era has started, the company’s own published definition, graded on the only peer-reviewed rubric, is asking for a system at the 99th percentile of skilled adults across a wide range of non-physical tasks. Not one at the 50th.

Here is what that grid looks like when you put the systems in it.

Where systems sit on the DeepMind Levels of AGI grid, September 2026A five-by-two grid. Rows are performance levels from Emerging at the bottom to Superhuman at the top. Columns are Narrow and General. The Narrow column is filled at every level: SHRDLU at Emerging, Siri and Alexa at Competent, Grammarly and Dall-E 2 at Expert, Deep Blue and AlphaGo at Exceptional, AlphaFold and Stockfish at Superhuman. The General column has ChatGPT, Bard, Llama 2 and Gemini at Emerging and nothing above it. GPT-6 Astra is marked at Competent General as a claim rather than a measurement, with no published assessment against the rubric.NarrowGeneralEmergingmatches an unskilled humanSHRDLU and rule-based systemsChatGPT, Bard, Llama 2, GeminiCompetent50th percentile, skilled adultsSiri, Alexa, Google AssistantGPT-6 Astra?claimed in a briefing, never assessedExpert90th percentileGrammarly, Dall-E 2no system placed hereExceptional99th percentileDeep Blue, AlphaGono system placed hereSuperhumanbeats every humanAlphaFold, Stockfishno system placed here
Where systems sit on the DeepMind Levels of AGI grid, September 2026
Performance levelCriterionNarrowGeneral
Superhumanbeats every humanAlphaFold, StockfishNo system placed here
Exceptional99th percentileDeep Blue, AlphaGoNo system placed here
Expert90th percentileGrammarly, Dall-E 2No system placed here
Competent50th percentile, skilled adultsSiri, Alexa, Google AssistantGPT-6 Astra?. claimed in a briefing, never assessed
Emergingmatches an unskilled humanSHRDLU and rule-based systemsChatGPT, Bard, Llama 2, Gemini
The Levels of AGI grid with the paper's own examples placed in it. Every rung of the Narrow column is occupied and has been for years. The General column is occupied on its bottom rung only. Competent General, the first rung where the word AGI starts to mean what people think it means, has never had a system placed in it, and nobody has published an assessment of Astra against that rubric.Source: Placements from Morris et al., Levels of AGI for Operationalizing Progress on the Path to AGI, arXiv preprint 2311.02462 v5. The GPT-6 Astra cell marks a claim made in a press briefing, not an assessment run against the paper's criteria

The empty region is the finding. Four cells in the General column have never held a system, and the argument about Astra is an argument about whether one of them should now be filled.

The paper is explicit about which one matters. Its own caption calls Competent AGI the level “which has not been achieved by any public systems at the time of writing” and says it “best corresponds to many prior conceptions of AGI.” That is the rung the word is doing work at.

The 10 tests, scored

1. ARC-AGI-3, the benchmark with AGI in its name

Partial. This is the strongest evidence in the set, and it is genuinely strong.

The ARC Prize Foundation’s writeup, published on launch day, reports that Astra beat the human baseline on action efficiency: it used fewer actions than the median tested human on 96.0% of levels, and 51.7% fewer actions per level on average. That is not a score. That is a claim about how quickly the model figures out an environment it has never seen, and it is the part that impressed the skeptics.

François Chollet, who designed the benchmark, described what Astra does as efficient on-the-fly symbolic world modelling, with the model inventing its own shorthand notation to represent each game state. Even Gary Marcus, who has spent a decade arguing that neural networks need symbolic world models bolted on, called it vindicating.

So why only partial? Because ARC Prize wrote the disclaimer themselves: saturating the benchmark “would not represent ‘proof of achieving AGI.’” Their reason is specific rather than modest. ARC-AGI-3 “has a tightly bounded scope and format, and its environments have deterministic, closed-ended mechanics and goals. It does not represent the complexity and open-endedness of the real world.”

Closed worlds with deterministic rules is the category where machines have been beating humans since Deep Blue.

2. Does the number survive a change of harness

Fail.

OpenAI’s headline was 98.6%, and the leaderboards now carry 99.9%. ARC Prize published two tables the same day. On its own standard harness, Astra scores 62.7% at max effort, for $26,098. Under OpenAI’s Provider Adapter harness, which preserves reasoning state between turns, it scores 99.9% at high effort, for $18,817.

Same model. Same benchmark. Same week. A 37-point spread produced by the scaffolding around the model rather than by the model.

The sibling analysis of what GPT-6 Astra actually costs works through the effort-level table in detail; the short version is that under the Provider Adapter every effort setting from none to max lands between 96.7% and 99.9%, which points at the adapter and not the reasoning.

A capability claim that swings 37 points on infrastructure is a claim about a system, not about a model. That distinction is the whole subject of how benchmark scores get gamed, and it applies here without modification.

Worth noting what the public leaderboards did with this. The llm-stats ARC-AGI-3 board, read on September 10, 2026, lists Astra at 0.999 above Claude Opus 5 at 0.302 and GPT-5.6 Sol at 0.078, and flags the whole table as five self-reported evaluations with zero verified results. The 99.9% is sitting on a ranked list next to numbers produced under different conditions.

3. Where it lands on the only published AGI rubric

No reading.

Nobody has run Astra against the Morris et al. criteria. Not OpenAI, not DeepMind, not an independent lab. The rubric exists, it is peer-reviewed, it is the most-cited attempt to make this question answerable, and in a week of AGI declarations not one party bothered to apply it.

The bar for Competent General is the 50th percentile of skilled adults across a wide range of non-physical tasks, including metacognitive ones like learning a new skill. Astra is plausibly there on some slices and obviously not on others. Absent an assessment, that is a guess, and a guess is what everyone making the claim is offering.

4. How long it can work unsupervised

No reading.

METR’s task-completion time horizon is the most-cited number in the autonomy debate: the length of task a model finishes with 50% reliability, which has been doubling roughly every four months.

METR has not measured Astra. Its time horizons page was last updated on May 8, 2026, and the newest model on it is GPT-5.4. A community analysis on LessWrong puts Astra’s 50% horizon somewhere between eight minutes and an hour, probably 15 to 40 minutes, but that is an estimate by a forecaster rather than a measurement by the lab that owns the metric.

The gap matters because of what the extrapolation says. Fit the doubling curve forward to September 2026 and you get tasks in the range of a day and a half of autonomous work. Estimate the actual model and you get well under an hour. Those two numbers cannot both be describing the same system, and only one of them is a measurement.

5. Can it do economically valuable work

Partial.

This is OpenAI’s own definition, so it is the fairest test to apply, and OpenAI built the benchmark for it. GDPval (Patwardhan et al., arXiv preprint October 2025, later an ICLR 2026 conference paper) covers 44 occupations across the nine sectors contributing most to US GDP, using real work products: a legal brief, an engineering blueprint, a nursing care plan. OpenAI reports that frontier models are approaching industry experts on deliverable quality, at roughly 100 times the speed and 100 times lower cost.

Approaching expert quality on graded deliverables is a real result and it is the closest anything comes to passing here. Two things hold it to partial. GDPval is the vendor’s own benchmark, graded on discrete work products rather than on holding down the job, and OpenAI’s charter says “most economically valuable work,” which is a statement about the labor market rather than about a task set. No labor-market series has moved in a way that would corroborate it.

Remember where the DeepMind researchers put that charter threshold: Exceptional AGI, the 99th percentile. Approaching the median expert on a graded task set is a long way from replacing the top one in a hundred at the job.

6. Do the independent aggregates agree

Fail. They do not even agree on the direction.

Epoch AI’s Capabilities Index puts Astra at 169 against GPT-5.6 Sol’s 162, and ranks it first overall. Artificial Analysis, publishing on September 9, 2026, scores Astra (max) at 53 on Intelligence Index v4.3, six points above Sol and level with Claude Fable 5.1. First place outright on one, a tie on the other.

Then there is what happened to the index itself. On September 4 this site recorded Astra at 61 and Fable 5.1 at 66 under Intelligence Index v4.1.1, a five-point deficit. Five days later, under v4.3, they are level. The models did not change. The ruler did.

GPT-6 Astra against GPT-5.6 Sol, four readings from three scoreboards, September 2026Four horizontal bars on a logarithmic axis showing GPT-6 Astra as a multiple of GPT-5.6 Sol. ARC-AGI-3 under the OpenAI Provider Adapter harness rises from 7.8 percent to 99.9 percent, a multiple of 12.8. ARC-AGI-3 under the ARC Prize standard harness rises from 7.8 percent to 62.7 percent, a multiple of 8.04. The Artificial Analysis Intelligence Index version 4.3 rises from 47 to 53, a multiple of 1.13. The Epoch Capabilities Index rises from 162 to 169, a multiple of 1.04.1x2x5x10x20xGPT-6 Astra as a multiple of GPT-5.6 Sol, log scale. 1x is GPT-5.6 Sol.ARC-AGI-3, provider adapter7.8% to 99.9%12.8xARC-AGI-3, standard harness7.8% to 62.7%8.04xAA Intelligence Index v4.347 to 531.13xEpoch Capabilities Index162 to 1691.04x
GPT-6 Astra against GPT-5.6 Sol, four readings from three scoreboards, September 2026
ScoreboardPrevious modelNew modelMultiple
ARC-AGI-3, provider adapter (7.8% to 99.9%)7.899.912.8x
ARC-AGI-3, standard harness (7.8% to 62.7%)7.862.78.04x
AA Intelligence Index v4.3 (47 to 53)47531.13x
Epoch Capabilities Index (162 to 169)1621691.04x
One model generation, three scoreboards, four readings, and no agreement about how big the jump was. On ARC Prize's own harness Astra is eight times its predecessor. On the aggregate indices it is a single-digit-percent improvement. The axis is logarithmic because a linear one would render the two aggregate rows as identical stubs, which is the disagreement being measured.Source: ARC-AGI-3 scores from the ARC Prize Foundation writeup and the llm-stats leaderboard; Epoch Capabilities Index as reported by The Decoder; Artificial Analysis Intelligence Index v4.3 from Artificial Analysis, with GPT-5.6 Sol at 47 derived from the stated six-point gain. All read September 10, 2026

Pick your scoreboard and the same six months of progress is either a 12.8-fold leap or a 4% improvement. Both readings are defensible. Neither is AGI-shaped on its own, and the spread between them is the reason which benchmarks deserve trust is a live question rather than a pedantic one.

7. Does it know when it does not know

Fail.

Artificial Analysis reports Astra’s hallucination rate falling from 92% to 51%. Read the improvement first: cutting confident wrongness almost in half between generations is a large engineering result, and it is the single number most likely to matter for anyone putting the model in a workflow.

Now read the level. On the eval that asks whether a model will assert something false rather than decline, it still does so about half the time.

Calibration is not a nice-to-have on the path to general intelligence. A system that cannot reliably represent the boundary of its own knowledge cannot be trusted to run unsupervised, which loops back to test 4. Whatever Astra is, it is not a thing you leave alone with a task and a credit card.

8. Do the people who measure this endorse the claim

Fail. Unanimously, and the list is not the usual suspects.

ARC Prize, whose benchmark produced the headline: “we are not claiming that it is AGI.” Chollet moved his personal AGI forecast forward from 2030, telling The Decoder “sooner, because progress is happening faster than I expected,” and still declined the label for Astra. Marcus, praising the symbolic world modelling in the same breath: success on ARC-AGI is impressive but “not, despite the name of the task, proof of AGI.” Yann LeCun continues to reject the framing, arguing that general intelligence is the wrong target and that systems still lack persistent world models and physically grounded planning, the position underlying the world-model funding wave.

The most telling absence is OpenAI’s. Its own model documentation for Astra makes no AGI claim at all. It calls the model “our most capable model, built for the hardest end-to-end work.” The AGI language lives in a press briefing and a social post, not in the model card, not in the system card, and not in anything a customer signs.

9. Have long real interest rates moved

Fail. This is the test almost nobody runs, and it is the most rigorous one available.

The argument comes from Transformative AI, existential risk, and real interest rates by Trevor Chow, Basil Halperin and J. Zachary Mazlish (working paper, drawing on 59 countries and 35 years of data). The logic is standard consumption smoothing. If transformative AI is close, future output is enormous relative to today, so rational agents borrow against it now, and long real rates rise. The conclusion holds even if you expect the technology to be catastrophic: the prospect of no future is also a reason to spend now rather than save.

Long real rates have not risen. Reporting the market reaction to Astra, Fortune noted that long-term Treasury yields have on average fallen by more than a tenth of a percentage point around major model releases. Bond markets are pricing the opposite of an intelligence explosion.

Halperin’s own summary, quoted in that piece: by rigorous standards, “we just absolutely have not achieved AGI, even though the models are astounding.”

10. Has the equity market repriced

Fail, and this is where the claim collides with the balance sheet.

Huang declared AGI on Sunday, September 6, 2026, crediting roughly 100,000 Nvidia Grace Blackwell NVL72 systems for training the model. Markets reopened Tuesday after Labor Day. Nvidia closed at $225.73, down 2.01%.

If a market genuinely believed general intelligence had arrived on hardware one company sells, that company’s shares would not slip 2%. Something did move: CoreWeave rose about 15% and SoftBank added roughly 2%, which reads as compute-supply repricing rather than an AGI repricing.

The obvious objection is that AGI is already priced in after three years of it being the explicit thesis. Fair. But that objection cuts against the announcement having information content, which is the same conclusion by a different route. Analyst Gil Luria offered a duller explanation to Fortune, that Nvidia is now too big to grow into news and investors treat it as a liquidity source.

What the tally means for the AI trade

Hyperscaler capital expenditure is running near $800 billion in 2026 with projections above $1.3 trillion for 2027. Those budgets do not need AGI. They need the current models to be useful enough to sell, which the evidence supports comfortably: Astra tops the practical scoreboards, and on cost per finished task it is the cheapest capable option on the model leaderboard.

The AGI framing does something specific to that trade, and it is worth being precise about who benefits. The declaration came from a model vendor’s president and a chip vendor’s CEO. It did not come from the benchmark authors, the independent evaluators, the academic rubric, or either market that would have to reprice if it were true. That is not a conspiracy. It is an incentive gradient, and it is legible.

The practical read for anyone allocating against this: nothing in Astra’s release changes a capex thesis, because nothing in it changes the unit economics that thesis rests on. The prices moved (Astra lists at $10 and $50 per million tokens against Sol’s $4 and $20) and the cost per solved task fell anyway. That is a normal, healthy, entirely non-apocalyptic generational improvement, and it is the thing worth underwriting.

Test 9 is the one to keep on a watchlist. Long real rates are the cleanest available signal, they are published daily, and they are not gameable by a harness change. When they start rising through a model release rather than falling, that is the day this scorecard gets interesting.

How this scorecard gets re-scored

These ten tests are model-agnostic on purpose. Every one of them is a bar somebody else published, so the next lab that declares AGI can be run through the identical rubric without inventing new criteria to fit the announcement.

Two tests currently return no reading. If METR publishes a time horizon for Astra, or anyone runs the Levels of AGI rubric against it, this post gets an updatedDate and the affected rows change. The tally at the top is a snapshot of published evidence as of September 10, 2026, not a permanent judgement.

Common questions

Did GPT-6 Astra really score 99.9% on ARC-AGI-3?

Under OpenAI’s Provider Adapter harness, yes. Under ARC Prize’s own standard harness the same model scores 62.7%. Both numbers were published by ARC Prize on the same day, and the 37-point gap comes from the scaffolding rather than the model.

Who has actually called GPT-6 Astra AGI?

Greg Brockman, OpenAI’s president, and Jensen Huang, Nvidia’s CEO. No benchmark author, independent evaluator or academic framework has. OpenAI’s own model documentation does not use the term.

What would change the verdict?

A published assessment against the Levels of AGI rubric showing Competent General or above, a METR time horizon in the range of days rather than minutes, and long-term real interest rates rising through a model release. Any one of those turns a row; all three together would flip the tally.

Sources

ARC Prize Foundation (2026). GPT-6 Astra on ARC-AGI-3. ARC Prize Foundation, independent benchmark. https://arcprize.org/blog/astra Verified 2026-09-10.

Morris, M.R., Sohl-Dickstein, J., Fiedel, N., Warkentin, T., Dafoe, A., Faust, A., Farabet, C., and Legg, S. (2023). Levels of AGI for Operationalizing Progress on the Path to AGI. arXiv preprint 2311.02462, v5; ICML 2024 position paper. https://arxiv.org/abs/2311.02462 Verified 2026-09-10.

Patwardhan, T., Dias, R., Proehl, E., et al. (2025). GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks. arXiv preprint 2510.04374; ICLR 2026 conference paper. https://arxiv.org/abs/2510.04374 Verified 2026-09-10.

Chow, T., Halperin, B., and Mazlish, J.Z. Transformative AI, existential risk, and real interest rates. Working paper. https://basilhalperin.com/papers/agi_emh.pdf Verified 2026-09-10.

METR (2026). Task-Completion Time Horizons of Frontier AI Models. METR, independent evaluation organisation. Page last updated 2026-05-08. https://metr.org/time-horizons/ Verified 2026-09-10.

Artificial Analysis (2026). Benchmarking GPT-6 Astra. Artificial Analysis, independent benchmarking service, Intelligence Index v4.3. https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra Verified 2026-09-10.

OpenAI (2026). GPT-6 Astra. OpenAI developer documentation, model reference. https://developers.openai.com/api/docs/models/gpt-6-astra Verified 2026-09-10.

llm-stats (2026). ARC-AGI-3 leaderboard. llm-stats, aggregator; entries flagged self-reported and unverified. https://llm-stats.com/benchmarks/arc-agi-3 Verified 2026-09-10.

The Decoder (2026). Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward. Secondary coverage; source of the Epoch Capabilities Index figures and the Chollet quote. https://the-decoder.com/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3-pulls-chollets-agi-forecast-forward/ Verified 2026-09-10.

Fortune (2026). AGI is here, Nvidia CEO Jensen Huang has declared. Here’s why the market doesn’t care. Secondary coverage; source of the Treasury-yield observation and the Halperin and Luria quotes. https://fortune.com/2026/09/09/markets-agi-nvidia-singularity-wall-street/ Verified 2026-09-10.

Marcus, G. (2026). Hot take on GPT-6 Astra. Marcus on AI, author’s own newsletter. https://garymarcus.substack.com/p/hot-take-on-gpt-6-astra Verified 2026-09-10.

Simon Willison (2026). Tracking the history of the now-deceased OpenAI Microsoft AGI clause. Secondary coverage of the AGI clause timeline. https://simonwillison.net/2026/Apr/27/now-deceased-agi-clause/ Verified 2026-09-10.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← Back to Models & benchmarks