Artificial Analysis Intelligence Index
Also known as AA Intelligence Index, AAII
The Artificial Analysis Intelligence Index is the single number most people mean when they say a model scores in the fifties or sixties. It is not a test. It is a weighted composite of about ten independent evaluations run by Artificial Analysis, an independent benchmarking firm, and its component set is revised often enough that the score is only meaningful when you say which version produced it.
| What it measures | A composite index of overall model intelligence aggregating performance across reasoning, coding, knowledge, science and agentic tasks. |
|---|---|
| Built by | Artificial Analysis (independent), 2024 |
| Format | A weighted aggregate of independent evaluations, revised often; v4.3 (September 2026) combines 11 evals in four categories: Agents 30%, General 30%, Coding 20%, Scientific Reasoning 20% |
| Scoring metric | Composite index score (0 to 100 aggregate); scores are not comparable across index versions |
| Status | Active |
| Representative top score | 53 (index v4.3) · Claude Fable 5.1 (max effort) and GPT-6 Astra (max) · read 2026-09 |
| Official leaderboard | artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index |
How Artificial Analysis Intelligence Index works
The index aggregates separate evaluations into four weighted categories. In v4.3, current as of September 2026, those are Agents at 30%, General at 30%, Coding at 20% and Scientific Reasoning at 20%. The components are AA-Briefcase at 15%, GDPval-AA v2 at 10% and AutomationBench-AA at 5% for agents; Terminal-Bench v4.0 and SciCode at 10% each for coding; AA-Omniscience accuracy at 10%, AA-Omniscience non-hallucination at 5%, GDP.pdf at 10% and AA-LCR v1.1 at 5% for general; and Humanity's Last Exam and CritPt at 10% each for scientific reasoning. Evaluations with private questions or answers carry 45% of the v4.3 weighting, up from 40% in v4.2, which is the firm's defence against labs training on the test.
History and current status
Artificial Analysis has run the index since 2024, revising it whenever components saturate. The recent cadence is the thing worth knowing. v4.1 shifted the index toward agentic workloads. v4.2 landed on 4 September 2026, adding AA-Briefcase and GDP.pdf, retiring GPQA Diamond as saturated, and doubling the private-test weighting to 40%. v4.3 landed on 7 September 2026, three days later, upgrading Terminal-Bench from v2.1 to v4.0, replacing t3-Banking with AutomationBench-AA, and raising private weighting to 45%. Two component changes in four days is not a stable yardstick.
What the score does not tell you
It is a vendor-maintained composite, so its absolute values are a property of the current weighting rather than of the models. Because components are swapped and reweighted, scores do not compare across versions, and nothing in a quoted number tells you which version it came from. The firm also publishes several reasoning-effort variants per model, max, xhigh, high, medium and low, which produce materially different scores for the same model name, so a citation without the effort setting is ambiguous. Per-evaluation columns have been dropped from the public leaderboard at times, which makes it harder to check what is driving a score. Many leaderboard values carry a trailing asterisk for which no public legend is published.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Artificial Analysis Intelligence Index is not one of the 15 benchmarks scored in the benchmark trust ranking, which rates the most-quoted tests on those same four axes and sorts them most trustworthy first.
What Artificial Analysis Intelligence Index scores actually mean
The clearest way to read this index is to watch one model across versions. Under v4.1 the top of the board ran to about 66. Under v4.2, checked on 5 September 2026, Claude Fable 5.1 topped it at 57. Under v4.3, checked on 12 September 2026, the top reads 53, shared by Claude Fable 5.1 at max and xhigh effort and GPT-6 Astra at max. The models did not get worse across those eight days; the index was recalibrated twice. So a gap of a few points between two models on the same version is a real if modest difference, while the same gap between two versions is noise introduced by the scoring. Treat the index as an ordering within one published version and never as an absolute level.
Who reports Artificial Analysis Intelligence Index, and how to read it
Almost everybody, usually without the version. Labs cite it in launch posts when it flatters them, aggregators such as Price Per Token draw on it downstream, and press coverage reduces it to a single figure. That downstream reach is precisely why the version matters: a number republished from a v4.1-era article and set beside a v4.3 number implies a decline that never happened. When you see an index score quoted, the three things to ask for are the version, the reasoning-effort variant, and the date it was read.
When to weight Artificial Analysis Intelligence Index in a model choice
Use it as a first-pass filter when you need one yardstick across many models, which is what it is good at, and then check the component that matches your actual workload. If you are buying coding capability, Terminal-Bench and SciCode carry 20% between them and the underlying boards are more informative than the composite. If you are buying agentic knowledge work, AA-Briefcase and GDPval-AA carry 25%. Record the version, the effort variant and the read date alongside any number you store, because an index figure without those three is not reproducible and will silently rot inside whatever you put it in.
Benchmarks to read alongside this one
Terminal-Bench
Whether an AI agent can complete hard, realistic command-line tasks (build, configure, train, debug, secure) end to end inside a real terminal.
GDPval
Whether a model can produce the actual deliverables of skilled professional work (documents, slides, spreadsheets, diagrams) well enough to stand against an industry expert's version.
Humanity's Last Exam
Frontier, closed-ended expert knowledge and reasoning across more than 100 academic disciplines at the limit of human expertise.
Artificial Analysis Intelligence Index: frequently asked questions
- What is the Artificial Analysis Intelligence Index?
- It is a weighted composite of about ten independent evaluations run by Artificial Analysis, reported as a single score. Version 4.3, current since 7 September 2026, weights Agents 30%, General 30%, Coding 20% and Scientific Reasoning 20%.
- Why did AA Intelligence Index scores drop in September 2026?
- Because the index was rebuilt twice in four days. v4.2 shipped on 4 September and v4.3 on 7 September. The top score went from about 66 under v4.1 to 57 under v4.2 to 53 under v4.3. That is recalibration, not models getting worse.
- Can I compare Intelligence Index scores across versions?
- No. Components are added, retired and reweighted between versions, so absolute values are only comparable within one version. Always record the version, the reasoning-effort variant such as max or high, and the date the score was read.
Sources
- Artificial Analysis (2026). Announcing the Artificial Analysis Intelligence Index v4.3, 7 September 2026. artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3
- Artificial Analysis (2026). Announcing Artificial Analysis Intelligence Index v4.2, 4 September 2026. artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-2
- Artificial Analysis. Intelligence benchmarking methodology. artificialanalysis.ai/methodology/intelligence-benchmarking
- Artificial Analysis. Artificial Analysis Intelligence Index leaderboard. artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.
← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →