Capital & Compute

METR Time Horizon

Benchmark· Human preference & holistic· Checked 2026-07-27

Also known as 50% task-completion time horizon

The METR time horizon is the only major capability measure expressed in units of human time rather than as a percentage. It reports the length of task, measured by how long people with relevant expertise take, that a model can complete with a 50% success rate. That framing makes it the single most quotable number in AI evaluation, because a duration is intuitively meaningful in a way that a benchmark percentage is not.

Key facts about the METR Time Horizon benchmark
What it measuresModel capability expressed in human time: the length of task, measured by how long humans take, that a model completes with 50% success.
Built byMETR (Kwa, West, Becker et al.), 2025
FormatTasks drawn from RE-Bench, HCAST and 66 shorter tasks, each timed with domain-expert humans to establish a human duration
Scoring metric50%-task-completion time horizon, in minutes or hours
StatusActive
Representative top scoreNot independently confirmed; see the leaderboard

How METR Time Horizon works

METR assembled tasks from RE-Bench, HCAST and 66 additional shorter tasks, then timed domain-expert humans completing each one to establish a human duration for it. Models are run on the same tasks, and the 50%-task-completion time horizon is the human task length at which model success falls to 50%. A model with a one-hour horizon can be expected to finish roughly half of the tasks that take a person about an hour.

History and current status

METR published the method in March 2025, reporting that frontier models of that period had a 50% time horizon of around 50 minutes and that the horizon had been doubling roughly every seven months since 2019. A 2026 update, Time Horizon 1.1, added tasks with longer human completion times and removed flawed ones, and the revised estimates put recent progress substantially faster than the original seven-month doubling.

What the score does not tell you

The measurement is confined to software and research-engineering tasks, so it is not a general statement about work. A 50% success rate is also a low bar for anything that matters: a model with a four-hour horizon fails half of four-hour tasks, which is not the same as being able to do them. And the horizon depends on the task distribution, which is why revising the task set revised the trend.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

Scored on those four axes, METR Time Horizon carries a concern score of 2 out of 8 and ranks 6 of 15, which puts it in the group worth quoting as it stands. See the reasoning behind that rating and how it compares with the rest of the field.

What METR Time Horizon scores actually mean

A time-horizon figure answers a different question from every percentage in this directory: how long a task, measured by how long it takes a domain-expert human, can a model complete with 50% reliability. That framing is why it is quotable and also why it is easy to misread. The 50% threshold is doing heavy lifting; the horizon at 80% reliability is substantially shorter, and 80% is closer to what unattended work requires. The launch paper reported the horizon doubling roughly every seven months since 2019, with later updates putting the recent rate higher. Because the estimate comes from fitting a curve across tasks of varying length, confidence intervals are wide and a single new data point can move the headline number noticeably.

Who reports METR Time Horizon, and how to read it

METR runs the evaluation itself and publishes the results, and the figure is widely cited in AI forecasting and policy discussion. Epoch AI’s Capabilities Index shows a comparable acceleration by an independent method, which is the best available corroboration. When citing a time horizon, state the version, because 1.0 and 1.1 are not the same measurement.

When to weight METR Time Horizon in a model choice

Use the time horizon for planning rather than for model selection: it is the best available answer to how much of a multi-hour job can be delegated today and how fast that is changing, which is a budgeting and staffing question more than a purchasing one. Always state the reliability threshold with the number, since a 50% horizon and an 80% horizon differ by a large factor and are frequently conflated in coverage. Treat the doubling rate as a trend estimate with real uncertainty, not a schedule. For a concrete build-or-buy decision, pair it with Terminal-Bench and its published run costs.

2025
First released
METR (Kwa, West, Becker et al.)
Active
Status today
As of September 12, 2026
See board
Representative top score
Not independently confirmed

Benchmarks to read alongside this one

METR Time Horizon: frequently asked questions

What is the METR time horizon?
It is a capability measure expressed in human time: the length of task, as timed with domain-expert humans, that a model completes with a 50% success rate. METR built it from RE-Bench, HCAST and 66 shorter tasks, publishing the method in March 2025.
How fast is the AI time horizon growing?
The original March 2025 paper reported a doubling roughly every seven months since 2019. The 2026 Time Horizon 1.1 update, which added longer tasks and removed flawed ones, put recent progress considerably faster than that original trend.
What does a 50% time horizon actually mean?
That the model succeeds on about half of the tasks of that human duration. It is a threshold, not a capability guarantee: a four-hour horizon means failing roughly half of four-hour tasks, which is why the number should not be read as "the AI can now do four hours of work."

Sources

  • Kwa, West, Becker, Deng et al. (2025). Measuring AI Ability to Complete Long Software Tasks. arXiv preprint 2503.14499. arxiv.org/abs/2503.14499
  • METR. Time-horizon research updates and methodology notes. metr.org

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →