Epoch Capabilities Index
Also known as ECI
The Epoch Capabilities Index is a single capability scale assembled from more than 50 separate benchmarks. Its purpose is to solve the problem that breaks every individual benchmark eventually: saturation. Because it bridges between benchmarks of different difficulty, the scale keeps working across periods in which any one component tops out.
| What it measures | Overall model capability on one continuous scale, stitched together from many benchmarks of differing difficulty. |
|---|---|
| Built by | Epoch AI, 2025 |
| Format | A latent-trait (item-response-theory) model over scores from more than 50 benchmarks, anchored so that Claude 3.5 Sonnet = 130 and GPT-5 = 150 |
| Scoring metric | ECI score on the anchored scale |
| Status | Active |
| Representative top score | Not independently confirmed; see the leaderboard |
| Official leaderboard | epoch.ai/eci |
How Epoch Capabilities Index works
ECI fits a one-dimensional latent-trait model, the same item-response-theory approach used in educational testing, over scores from more than 50 benchmarks. Relative benchmark difficulty is inferred wherever models have been evaluated on more than one benchmark, which is what lets the components be stitched together. The scale is anchored by two fixed points, Claude 3.5 Sonnet at 130 and GPT-5 at 150, and harder benchmarks carry more weight, so a model gains more from progress on difficult evaluations.
History and current status
Epoch AI introduced the index in 2025 and has maintained it since, expanding the component set as new evaluations appear. Its most cited output is a finding about the rate of progress: frontier improvement roughly doubled in pace, from about 8 points per year before April 2024 to about 15 points per year afterwards, coinciding with the arrival of reasoning models. The methodology is documented in the paper A Rosetta Stone for AI Benchmarks.
What the score does not tell you
Compressing many benchmarks into one number necessarily discards the shape of a model’s ability, and a single scalar invites exactly the over-reading that benchmark critics warn about. The index also depends on which benchmarks are included and which models were evaluated on what, so the scale reflects the coverage of the underlying data. The anchoring choice is a convention, not a measurement.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Epoch Capabilities Index is not one of the 15 benchmarks scored in the benchmark trust ranking, which rates the most-quoted tests on those same four axes and sorts them most trustworthy first.
What Epoch Capabilities Index scores actually mean
ECI is an anchored scale, not a percentage, and reading it requires the anchors: Claude 3.5 Sonnet sits at 130 and GPT-5 at 150. Because it fits an item-response-theory model across more than 50 benchmarks, weighting harder ones more heavily, a difference of a few points reflects a real shift in estimated capability rather than a rounding artifact on one saturated test. The trade-off is indirectness. ECI inherits whatever biases exist in its constituent benchmarks, and because those benchmarks are added and retired over time, the index is periodically recomputed, so a value quoted months ago may not match today's figure for the same model.
Who reports Epoch Capabilities Index, and how to read it
Epoch AI is an independent research organisation rather than a model vendor, which makes ECI one of the few composite indices not run by an interested party. Note that the underlying methodology paper was funded by Google DeepMind and written with researchers from its safety team, which is worth stating when citing it.
When to weight Epoch Capabilities Index in a model choice
Use ECI when you want one defensible number for how capable a model is in general, which is the question every individual benchmark has stopped answering as it saturates. It is the right citation for a trend argument about capability over time, and the wrong one for a specific task decision, where a targeted benchmark will always tell you more. Prefer it over any single-benchmark composite maintained by a model vendor, since Epoch is an independent research group with published methodology. Record the read date with the score, because the index is recomputed as its constituent benchmarks change.
Benchmarks to read alongside this one
METR Time Horizon
Model capability expressed in human time: the length of task, measured by how long humans take, that a model completes with 50% success.
Artificial Analysis Intelligence Index
A composite index of overall model intelligence aggregating performance across reasoning, coding, knowledge, science and agentic tasks.
LiveBench
Broad capability across six categories at once (math, coding, reasoning, data analysis, instruction following, language), on questions refreshed monthly.
Epoch Capabilities Index: frequently asked questions
- What is the Epoch Capabilities Index?
- ECI is a composite metric from Epoch AI that combines scores from more than 50 benchmarks into one general capability scale using an item-response-theory model. It is anchored so that Claude 3.5 Sonnet sits at 130 and GPT-5 at 150.
- Why does a composite index help with benchmark saturation?
- Because individual benchmarks stop discriminating once models reach their ceiling, but a composite that bridges benchmarks of different difficulty can keep measuring progress by shifting weight onto harder components. That lets one scale span years in which several components saturate.
- Is the Epoch Capabilities Index independent?
- Epoch AI is an independent research organisation, not a model vendor, which is unusual for a composite index. The underlying methodology paper, A Rosetta Stone for AI Benchmarks, was funded by Google DeepMind and co-written with its safety researchers.
Sources
- Epoch AI. Epoch Capabilities Index documentation, including the item-response-theory methodology and anchoring. epoch.ai/data/eci-documentation
- Epoch AI. Live ECI scores by model. epoch.ai/eci
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.
← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →