Epoch Capabilities Index
Also known as ECI
The Epoch Capabilities Index is a single capability scale assembled from more than 50 separate benchmarks. Its purpose is to solve the problem that breaks every individual benchmark eventually: saturation. Because it bridges between benchmarks of different difficulty, the scale keeps working across periods in which any one component tops out.
| What it measures | Overall model capability on one continuous scale, stitched together from many benchmarks of differing difficulty. |
|---|---|
| Built by | Epoch AI, 2025 |
| Format | A latent-trait (item-response-theory) model over scores from more than 50 benchmarks, anchored so that Claude 3.5 Sonnet = 130 and GPT-5 = 150 |
| Scoring metric | ECI score on the anchored scale |
| Status | Active |
| Representative top score | Not independently confirmed; see the leaderboard |
| Official leaderboard | epoch.ai/eci |
How Epoch Capabilities Index works
ECI fits a one-dimensional latent-trait model, the same item-response-theory approach used in educational testing, over scores from more than 50 benchmarks. Relative benchmark difficulty is inferred wherever models have been evaluated on more than one benchmark, which is what lets the components be stitched together. The scale is anchored by two fixed points, Claude 3.5 Sonnet at 130 and GPT-5 at 150, and harder benchmarks carry more weight, so a model gains more from progress on difficult evaluations.
History and current status
Epoch AI introduced the index in 2025 and has maintained it since, expanding the component set as new evaluations appear. Its most cited output is a finding about the rate of progress: frontier improvement roughly doubled in pace, from about 8 points per year before April 2024 to about 15 points per year afterwards, coinciding with the arrival of reasoning models. The methodology is documented in the paper A Rosetta Stone for AI Benchmarks.
What the score does not tell you
Compressing many benchmarks into one number necessarily discards the shape of a model’s ability, and a single scalar invites exactly the over-reading that benchmark critics warn about. The index also depends on which benchmarks are included and which models were evaluated on what, so the scale reflects the coverage of the underlying data. The anchoring choice is a convention, not a measurement.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Who reports Epoch Capabilities Index, and how to read it
Epoch AI is an independent research organisation rather than a model vendor, which makes ECI one of the few composite indices not run by an interested party. Note that the underlying methodology paper was funded by Google DeepMind and written with researchers from its safety team, which is worth stating when citing it.
Benchmarks to read alongside this one
METR Time Horizon
Model capability expressed in human time: the length of task, measured by how long humans take, that a model completes with 50% success.
Artificial Analysis Intelligence Index
A composite index of overall model intelligence aggregating performance across reasoning, coding, knowledge, science and agentic tasks.
LiveBench
Broad capability across six categories at once (math, coding, reasoning, data analysis, instruction following, language), on questions refreshed monthly.
Epoch Capabilities Index: frequently asked questions
- What is the Epoch Capabilities Index?
- ECI is a composite metric from Epoch AI that combines scores from more than 50 benchmarks into one general capability scale using an item-response-theory model. It is anchored so that Claude 3.5 Sonnet sits at 130 and GPT-5 at 150.
- Why does a composite index help with benchmark saturation?
- Because individual benchmarks stop discriminating once models reach their ceiling, but a composite that bridges benchmarks of different difficulty can keep measuring progress by shifting weight onto harder components. That lets one scale span years in which several components saturate.
- Is the Epoch Capabilities Index independent?
- Epoch AI is an independent research organisation, not a model vendor, which is unusual for a composite index. The underlying methodology paper, A Rosetta Stone for AI Benchmarks, was funded by Google DeepMind and co-written with its safety researchers.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.