METR Time Horizon
Also known as 50% task-completion time horizon
The METR time horizon is the only major capability measure expressed in units of human time rather than as a percentage. It reports the length of task, measured by how long people with relevant expertise take, that a model can complete with a 50% success rate. That framing makes it the single most quotable number in AI evaluation, because a duration is intuitively meaningful in a way that a benchmark percentage is not.
| What it measures | Model capability expressed in human time: the length of task, measured by how long humans take, that a model completes with 50% success. |
|---|---|
| Built by | METR (Kwa, West, Becker et al.), 2025 |
| Format | Tasks drawn from RE-Bench, HCAST and 66 shorter tasks, each timed with domain-expert humans to establish a human duration |
| Scoring metric | 50%-task-completion time horizon, in minutes or hours |
| Status | Active |
| Representative top score | Not independently confirmed; see the leaderboard |
How METR Time Horizon works
METR assembled tasks from RE-Bench, HCAST and 66 additional shorter tasks, then timed domain-expert humans completing each one to establish a human duration for it. Models are run on the same tasks, and the 50%-task-completion time horizon is the human task length at which model success falls to 50%. A model with a one-hour horizon can be expected to finish roughly half of the tasks that take a person about an hour.
History and current status
METR published the method in March 2025, reporting that frontier models of that period had a 50% time horizon of around 50 minutes and that the horizon had been doubling roughly every seven months since 2019. A 2026 update, Time Horizon 1.1, added tasks with longer human completion times and removed flawed ones, and the revised estimates put recent progress substantially faster than the original seven-month doubling.
What the score does not tell you
The measurement is confined to software and research-engineering tasks, so it is not a general statement about work. A 50% success rate is also a low bar for anything that matters: a model with a four-hour horizon fails half of four-hour tasks, which is not the same as being able to do them. And the horizon depends on the task distribution, which is why revising the task set revised the trend.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Who reports METR Time Horizon, and how to read it
METR runs the evaluation itself and publishes the results, and the figure is widely cited in AI forecasting and policy discussion. Epoch AI’s Capabilities Index shows a comparable acceleration by an independent method, which is the best available corroboration. When citing a time horizon, state the version, because 1.0 and 1.1 are not the same measurement.
Benchmarks to read alongside this one
Epoch Capabilities Index
Overall model capability on one continuous scale, stitched together from many benchmarks of differing difficulty.
GDPval
Whether a model can produce the actual deliverables of skilled professional work (documents, slides, spreadsheets, diagrams) well enough to stand against an industry expert's version.
TheAgentCompany
Whether an agent can do real knowledge work inside a simulated software company: browsing, coding, using internal tools, and messaging simulated colleagues.
METR Time Horizon: frequently asked questions
- What is the METR time horizon?
- It is a capability measure expressed in human time: the length of task, as timed with domain-expert humans, that a model completes with a 50% success rate. METR built it from RE-Bench, HCAST and 66 shorter tasks, publishing the method in March 2025.
- How fast is the AI time horizon growing?
- The original March 2025 paper reported a doubling roughly every seven months since 2019. The 2026 Time Horizon 1.1 update, which added longer tasks and removed flawed ones, put recent progress considerably faster than that original trend.
- What does a 50% time horizon actually mean?
- That the model succeeds on about half of the tasks of that human duration. It is a threshold, not a capability guarantee: a four-hour horizon means failing roughly half of four-hour tasks, which is why the number should not be read as "the AI can now do four hours of work."
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.