Is GPT-6 Astra AGI? The 10 Tests That Decide It
OpenAI says the AGI era has started. Ten published tests, from ARC-AGI-3 to the market reaction, scored against the evidence with an honest tally.
No. Not on any test anyone has actually published a result for.
OpenAI released GPT-6 Astra on September 3, 2026, and president Greg Brockman closed the briefing with “Welcome to the AGI era.” Three days later Nvidia’s Jensen Huang posted that AGI had arrived. Below are ten tests that could settle the question, each one scored against a bar somebody else set: the benchmark authors, a peer-reviewed taxonomy, OpenAI’s own charter, an independent evaluator, the bond market. Astra passes none of them outright. It comes close on two.
| # | Test | Bar set by | Verdict |
|---|---|---|---|
| 1 | ARC-AGI-3 saturation | ARC Prize Foundation | Partial |
| 2 | Does the score survive a harness change | ARC Prize Foundation | Fail |
| 3 | Placement on the Levels of AGI ladder | Morris et al., Google DeepMind | No reading |
| 4 | Autonomous task length | METR | No reading |
| 5 | Economically valuable work | OpenAI’s own charter | Partial |
| 6 | Do the independent aggregates agree | Artificial Analysis, Epoch AI | Fail |
| 7 | Does it know when it does not know | Artificial Analysis | Fail |
| 8 | Do the measurers endorse the claim | ARC Prize, Chollet, Marcus, LeCun | Fail |
| 9 | Have long real interest rates moved | Chow, Halperin and Mazlish | Fail |
| 10 | Has the equity market repriced | The tape | Fail |
Two partials, two tests with no published reading at all, six failures, zero passes.
That tally is not a verdict on the model. Astra is the strongest system anyone has shipped, and on one measure it is genuinely startling. It is a verdict on the claim.
What would count as AGI, and who gets to say
The reason “is it AGI” produces such bad arguments is that the term has three incompatible definitions in active use, and the people making claims rarely say which one they mean.
OpenAI’s founding definition is economic: highly autonomous systems that outperform humans at most economically valuable work. That is a labor-market statement, testable in principle and nowhere near satisfied.
The second definition was financial, and it is dead. Microsoft and OpenAI’s contract at one point pegged AGI to a profit threshold reported at $100 billion, which meant the most consequential AGI declaration in the industry was a revenue milestone. That clause mutated from a capability test in 2018, to the profit number in December 2024, to a procedural sunset, and was removed entirely in April 2026 when Microsoft took a non-exclusive license through 2032. Nobody has replaced it with anything.
The third is the only one with a published rubric behind it. In Levels of AGI for Operationalizing Progress on the Path to AGI, Meredith Ringel Morris, Jascha Sohl-Dickstein, Shane Legg and colleagues at Google DeepMind (arXiv preprint, November 2023; ICML 2024 position paper) split the question into two axes that people constantly conflate. Performance is depth: how good, measured against percentiles of skilled adults. Generality is breadth: how many kinds of task. A chess engine is Superhuman on one axis and Narrow on the other, and calling it intelligent or not intelligent depends entirely on which axis you were looking at.
Their ladder runs Emerging (a match for an unskilled human), Competent (50th percentile of skilled adults), Expert (90th), Exceptional (99th), Superhuman (better than every human). Version 5 of the paper renamed the fourth rung from Virtuoso to Exceptional, which is worth knowing if you are comparing against older write-ups that use the old word.
The useful part is that Morris and colleagues placed the rival definitions on their own ladder. Aguera y Arcas and Norvig’s definition, they write, “would fall into the ‘Emerging AGI’ category of our ontology,” which is a bar frontier models cleared years ago. And OpenAI’s own threshold of labor replacement “better matches ‘Exceptional AGI’”: the 99th percentile of skilled adults, two full rungs above the level nobody has reached yet.
So when OpenAI’s president says the AGI era has started, the company’s own published definition, graded on the only peer-reviewed rubric, is asking for a system at the 99th percentile of skilled adults across a wide range of non-physical tasks. Not one at the 50th.
Here is what that grid looks like when you put the systems in it.
| Performance level | Criterion | Narrow | General |
|---|---|---|---|
| Superhuman | beats every human | AlphaFold, Stockfish | No system placed here |
| Exceptional | 99th percentile | Deep Blue, AlphaGo | No system placed here |
| Expert | 90th percentile | Grammarly, Dall-E 2 | No system placed here |
| Competent | 50th percentile, skilled adults | Siri, Alexa, Google Assistant | GPT-6 Astra?. claimed in a briefing, never assessed |
| Emerging | matches an unskilled human | SHRDLU and rule-based systems | ChatGPT, Bard, Llama 2, Gemini |
The empty region is the finding. Four cells in the General column have never held a system, and the argument about Astra is an argument about whether one of them should now be filled.
The paper is explicit about which one matters. Its own caption calls Competent AGI the level “which has not been achieved by any public systems at the time of writing” and says it “best corresponds to many prior conceptions of AGI.” That is the rung the word is doing work at.
The 10 tests, scored
1. ARC-AGI-3, the benchmark with AGI in its name
Partial. This is the strongest evidence in the set, and it is genuinely strong.
The ARC Prize Foundation’s writeup, published on launch day, reports that Astra beat the human baseline on action efficiency: it used fewer actions than the median tested human on 96.0% of levels, and 51.7% fewer actions per level on average. That is not a score. That is a claim about how quickly the model figures out an environment it has never seen, and it is the part that impressed the skeptics.
François Chollet, who designed the benchmark, described what Astra does as efficient on-the-fly symbolic world modelling, with the model inventing its own shorthand notation to represent each game state. Even Gary Marcus, who has spent a decade arguing that neural networks need symbolic world models bolted on, called it vindicating.
So why only partial? Because ARC Prize wrote the disclaimer themselves: saturating the benchmark “would not represent ‘proof of achieving AGI.’” Their reason is specific rather than modest. ARC-AGI-3 “has a tightly bounded scope and format, and its environments have deterministic, closed-ended mechanics and goals. It does not represent the complexity and open-endedness of the real world.”
Closed worlds with deterministic rules is the category where machines have been beating humans since Deep Blue.
2. Does the number survive a change of harness
Fail.
OpenAI’s headline was 98.6%, and the leaderboards now carry 99.9%. ARC Prize published two tables the same day. On its own standard harness, Astra scores 62.7% at max effort, for $26,098. Under OpenAI’s Provider Adapter harness, which preserves reasoning state between turns, it scores 99.9% at high effort, for $18,817.
Same model. Same benchmark. Same week. A 37-point spread produced by the scaffolding around the model rather than by the model.
The sibling analysis of what GPT-6 Astra actually costs works through the effort-level table in detail; the short version is that under the Provider Adapter every effort setting from none to max lands between 96.7% and 99.9%, which points at the adapter and not the reasoning.
A capability claim that swings 37 points on infrastructure is a claim about a system, not about a model. That distinction is the whole subject of how benchmark scores get gamed, and it applies here without modification.
Worth noting what the public leaderboards did with this. The llm-stats ARC-AGI-3 board, read on September 10, 2026, lists Astra at 0.999 above Claude Opus 5 at 0.302 and GPT-5.6 Sol at 0.078, and flags the whole table as five self-reported evaluations with zero verified results. The 99.9% is sitting on a ranked list next to numbers produced under different conditions.
3. Where it lands on the only published AGI rubric
No reading.
Nobody has run Astra against the Morris et al. criteria. Not OpenAI, not DeepMind, not an independent lab. The rubric exists, it is peer-reviewed, it is the most-cited attempt to make this question answerable, and in a week of AGI declarations not one party bothered to apply it.
The bar for Competent General is the 50th percentile of skilled adults across a wide range of non-physical tasks, including metacognitive ones like learning a new skill. Astra is plausibly there on some slices and obviously not on others. Absent an assessment, that is a guess, and a guess is what everyone making the claim is offering.
4. How long it can work unsupervised
No reading.
METR’s task-completion time horizon is the most-cited number in the autonomy debate: the length of task a model finishes with 50% reliability, which has been doubling roughly every four months.
METR has not measured Astra. Its time horizons page was last updated on May 8, 2026, and the newest model on it is GPT-5.4. A community analysis on LessWrong puts Astra’s 50% horizon somewhere between eight minutes and an hour, probably 15 to 40 minutes, but that is an estimate by a forecaster rather than a measurement by the lab that owns the metric.
The gap matters because of what the extrapolation says. Fit the doubling curve forward to September 2026 and you get tasks in the range of a day and a half of autonomous work. Estimate the actual model and you get well under an hour. Those two numbers cannot both be describing the same system, and only one of them is a measurement.
5. Can it do economically valuable work
Partial.
This is OpenAI’s own definition, so it is the fairest test to apply, and OpenAI built the benchmark for it. GDPval (Patwardhan et al., arXiv preprint October 2025, later an ICLR 2026 conference paper) covers 44 occupations across the nine sectors contributing most to US GDP, using real work products: a legal brief, an engineering blueprint, a nursing care plan. OpenAI reports that frontier models are approaching industry experts on deliverable quality, at roughly 100 times the speed and 100 times lower cost.
Approaching expert quality on graded deliverables is a real result and it is the closest anything comes to passing here. Two things hold it to partial. GDPval is the vendor’s own benchmark, graded on discrete work products rather than on holding down the job, and OpenAI’s charter says “most economically valuable work,” which is a statement about the labor market rather than about a task set. No labor-market series has moved in a way that would corroborate it.
Remember where the DeepMind researchers put that charter threshold: Exceptional AGI, the 99th percentile. Approaching the median expert on a graded task set is a long way from replacing the top one in a hundred at the job.
6. Do the independent aggregates agree
Fail. They do not even agree on the direction.
Epoch AI’s Capabilities Index puts Astra at 169 against GPT-5.6 Sol’s 162, and ranks it first overall. Artificial Analysis, publishing on September 9, 2026, scores Astra (max) at 53 on Intelligence Index v4.3, six points above Sol and level with Claude Fable 5.1. First place outright on one, a tie on the other.
Then there is what happened to the index itself. On September 4 this site recorded Astra at 61 and Fable 5.1 at 66 under Intelligence Index v4.1.1, a five-point deficit. Five days later, under v4.3, they are level. The models did not change. The ruler did.
| Scoreboard | Previous model | New model | Multiple |
|---|---|---|---|
| ARC-AGI-3, provider adapter (7.8% to 99.9%) | 7.8 | 99.9 | 12.8x |
| ARC-AGI-3, standard harness (7.8% to 62.7%) | 7.8 | 62.7 | 8.04x |
| AA Intelligence Index v4.3 (47 to 53) | 47 | 53 | 1.13x |
| Epoch Capabilities Index (162 to 169) | 162 | 169 | 1.04x |
Pick your scoreboard and the same six months of progress is either a 12.8-fold leap or a 4% improvement. Both readings are defensible. Neither is AGI-shaped on its own, and the spread between them is the reason which benchmarks deserve trust is a live question rather than a pedantic one.
7. Does it know when it does not know
Fail.
Artificial Analysis reports Astra’s hallucination rate falling from 92% to 51%. Read the improvement first: cutting confident wrongness almost in half between generations is a large engineering result, and it is the single number most likely to matter for anyone putting the model in a workflow.
Now read the level. On the eval that asks whether a model will assert something false rather than decline, it still does so about half the time.
Calibration is not a nice-to-have on the path to general intelligence. A system that cannot reliably represent the boundary of its own knowledge cannot be trusted to run unsupervised, which loops back to test 4. Whatever Astra is, it is not a thing you leave alone with a task and a credit card.
8. Do the people who measure this endorse the claim
Fail. Unanimously, and the list is not the usual suspects.
ARC Prize, whose benchmark produced the headline: “we are not claiming that it is AGI.” Chollet moved his personal AGI forecast forward from 2030, telling The Decoder “sooner, because progress is happening faster than I expected,” and still declined the label for Astra. Marcus, praising the symbolic world modelling in the same breath: success on ARC-AGI is impressive but “not, despite the name of the task, proof of AGI.” Yann LeCun continues to reject the framing, arguing that general intelligence is the wrong target and that systems still lack persistent world models and physically grounded planning, the position underlying the world-model funding wave.
The most telling absence is OpenAI’s. Its own model documentation for Astra makes no AGI claim at all. It calls the model “our most capable model, built for the hardest end-to-end work.” The AGI language lives in a press briefing and a social post, not in the model card, not in the system card, and not in anything a customer signs.
9. Have long real interest rates moved
Fail. This is the test almost nobody runs, and it is the most rigorous one available.
The argument comes from Transformative AI, existential risk, and real interest rates by Trevor Chow, Basil Halperin and J. Zachary Mazlish (working paper, drawing on 59 countries and 35 years of data). The logic is standard consumption smoothing. If transformative AI is close, future output is enormous relative to today, so rational agents borrow against it now, and long real rates rise. The conclusion holds even if you expect the technology to be catastrophic: the prospect of no future is also a reason to spend now rather than save.
Long real rates have not risen. Reporting the market reaction to Astra, Fortune noted that long-term Treasury yields have on average fallen by more than a tenth of a percentage point around major model releases. Bond markets are pricing the opposite of an intelligence explosion.
Halperin’s own summary, quoted in that piece: by rigorous standards, “we just absolutely have not achieved AGI, even though the models are astounding.”
10. Has the equity market repriced
Fail, and this is where the claim collides with the balance sheet.
Huang declared AGI on Sunday, September 6, 2026, crediting roughly 100,000 Nvidia Grace Blackwell NVL72 systems for training the model. Markets reopened Tuesday after Labor Day. Nvidia closed at $225.73, down 2.01%.
If a market genuinely believed general intelligence had arrived on hardware one company sells, that company’s shares would not slip 2%. Something did move: CoreWeave rose about 15% and SoftBank added roughly 2%, which reads as compute-supply repricing rather than an AGI repricing.
The obvious objection is that AGI is already priced in after three years of it being the explicit thesis. Fair. But that objection cuts against the announcement having information content, which is the same conclusion by a different route. Analyst Gil Luria offered a duller explanation to Fortune, that Nvidia is now too big to grow into news and investors treat it as a liquidity source.
What the tally means for the AI trade
Hyperscaler capital expenditure is running near $800 billion in 2026 with projections above $1.3 trillion for 2027. Those budgets do not need AGI. They need the current models to be useful enough to sell, which the evidence supports comfortably: Astra tops the practical scoreboards, and on cost per finished task it is the cheapest capable option on the model leaderboard.
The AGI framing does something specific to that trade, and it is worth being precise about who benefits. The declaration came from a model vendor’s president and a chip vendor’s CEO. It did not come from the benchmark authors, the independent evaluators, the academic rubric, or either market that would have to reprice if it were true. That is not a conspiracy. It is an incentive gradient, and it is legible.
The practical read for anyone allocating against this: nothing in Astra’s release changes a capex thesis, because nothing in it changes the unit economics that thesis rests on. The prices moved (Astra lists at $10 and $50 per million tokens against Sol’s $4 and $20) and the cost per solved task fell anyway. That is a normal, healthy, entirely non-apocalyptic generational improvement, and it is the thing worth underwriting.
Test 9 is the one to keep on a watchlist. Long real rates are the cleanest available signal, they are published daily, and they are not gameable by a harness change. When they start rising through a model release rather than falling, that is the day this scorecard gets interesting.
How this scorecard gets re-scored
These ten tests are model-agnostic on purpose. Every one of them is a bar somebody else published, so the next lab that declares AGI can be run through the identical rubric without inventing new criteria to fit the announcement.
Two tests currently return no reading. If METR publishes a time horizon for Astra, or anyone runs the Levels of AGI rubric against it, this post gets an updatedDate and the affected rows change. The tally at the top is a snapshot of published evidence as of September 10, 2026, not a permanent judgement.
Common questions
Did GPT-6 Astra really score 99.9% on ARC-AGI-3?
Under OpenAI’s Provider Adapter harness, yes. Under ARC Prize’s own standard harness the same model scores 62.7%. Both numbers were published by ARC Prize on the same day, and the 37-point gap comes from the scaffolding rather than the model.
Who has actually called GPT-6 Astra AGI?
Greg Brockman, OpenAI’s president, and Jensen Huang, Nvidia’s CEO. No benchmark author, independent evaluator or academic framework has. OpenAI’s own model documentation does not use the term.
What would change the verdict?
A published assessment against the Levels of AGI rubric showing Competent General or above, a METR time horizon in the range of days rather than minutes, and long-term real interest rates rising through a model release. Any one of those turns a row; all three together would flip the tally.
Sources
ARC Prize Foundation (2026). GPT-6 Astra on ARC-AGI-3. ARC Prize Foundation, independent benchmark. https://arcprize.org/blog/astra Verified 2026-09-10.
Morris, M.R., Sohl-Dickstein, J., Fiedel, N., Warkentin, T., Dafoe, A., Faust, A., Farabet, C., and Legg, S. (2023). Levels of AGI for Operationalizing Progress on the Path to AGI. arXiv preprint 2311.02462, v5; ICML 2024 position paper. https://arxiv.org/abs/2311.02462 Verified 2026-09-10.
Patwardhan, T., Dias, R., Proehl, E., et al. (2025). GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks. arXiv preprint 2510.04374; ICLR 2026 conference paper. https://arxiv.org/abs/2510.04374 Verified 2026-09-10.
Chow, T., Halperin, B., and Mazlish, J.Z. Transformative AI, existential risk, and real interest rates. Working paper. https://basilhalperin.com/papers/agi_emh.pdf Verified 2026-09-10.
METR (2026). Task-Completion Time Horizons of Frontier AI Models. METR, independent evaluation organisation. Page last updated 2026-05-08. https://metr.org/time-horizons/ Verified 2026-09-10.
Artificial Analysis (2026). Benchmarking GPT-6 Astra. Artificial Analysis, independent benchmarking service, Intelligence Index v4.3. https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra Verified 2026-09-10.
OpenAI (2026). GPT-6 Astra. OpenAI developer documentation, model reference. https://developers.openai.com/api/docs/models/gpt-6-astra Verified 2026-09-10.
llm-stats (2026). ARC-AGI-3 leaderboard. llm-stats, aggregator; entries flagged self-reported and unverified. https://llm-stats.com/benchmarks/arc-agi-3 Verified 2026-09-10.
The Decoder (2026). Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward. Secondary coverage; source of the Epoch Capabilities Index figures and the Chollet quote. https://the-decoder.com/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3-pulls-chollets-agi-forecast-forward/ Verified 2026-09-10.
Fortune (2026). AGI is here, Nvidia CEO Jensen Huang has declared. Here’s why the market doesn’t care. Secondary coverage; source of the Treasury-yield observation and the Halperin and Luria quotes. https://fortune.com/2026/09/09/markets-agi-nvidia-singularity-wall-street/ Verified 2026-09-10.
Marcus, G. (2026). Hot take on GPT-6 Astra. Marcus on AI, author’s own newsletter. https://garymarcus.substack.com/p/hot-take-on-gpt-6-astra Verified 2026-09-10.
Simon Willison (2026). Tracking the history of the now-deceased OpenAI Microsoft AGI clause. Secondary coverage of the AGI clause timeline. https://simonwillison.net/2026/Apr/27/now-deceased-agi-clause/ Verified 2026-09-10.