Capital & Compute
Benchmark· Human preference & holistic· Checked 2026-07-27

METR Time Horizon

Also known as 50% task-completion time horizon

The METR time horizon is the only major capability measure expressed in units of human time rather than as a percentage. It reports the length of task, measured by how long people with relevant expertise take, that a model can complete with a 50% success rate. That framing makes it the single most quotable number in AI evaluation, because a duration is intuitively meaningful in a way that a benchmark percentage is not.

Key facts about the METR Time Horizon benchmark
What it measuresModel capability expressed in human time: the length of task, measured by how long humans take, that a model completes with 50% success.
Built byMETR (Kwa, West, Becker et al.), 2025
FormatTasks drawn from RE-Bench, HCAST and 66 shorter tasks, each timed with domain-expert humans to establish a human duration
Scoring metric50%-task-completion time horizon, in minutes or hours
StatusActive
Representative top scoreNot independently confirmed; see the leaderboard

How METR Time Horizon works

METR assembled tasks from RE-Bench, HCAST and 66 additional shorter tasks, then timed domain-expert humans completing each one to establish a human duration for it. Models are run on the same tasks, and the 50%-task-completion time horizon is the human task length at which model success falls to 50%. A model with a one-hour horizon can be expected to finish roughly half of the tasks that take a person about an hour.

History and current status

METR published the method in March 2025, reporting that frontier models of that period had a 50% time horizon of around 50 minutes and that the horizon had been doubling roughly every seven months since 2019. A 2026 update, Time Horizon 1.1, added tasks with longer human completion times and removed flawed ones, and the revised estimates put recent progress substantially faster than the original seven-month doubling.

What the score does not tell you

The measurement is confined to software and research-engineering tasks, so it is not a general statement about work. A 50% success rate is also a low bar for anything that matters: a model with a four-hour horizon fails half of four-hour tasks, which is not the same as being able to do them. And the horizon depends on the task distribution, which is why revising the task set revised the trend.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

Who reports METR Time Horizon, and how to read it

METR runs the evaluation itself and publishes the results, and the figure is widely cited in AI forecasting and policy discussion. Epoch AI’s Capabilities Index shows a comparable acceleration by an independent method, which is the best available corroboration. When citing a time horizon, state the version, because 1.0 and 1.1 are not the same measurement.

2025
First released
METR (Kwa, West, Becker et al.)
Active
Status today
As of July 27, 2026
See board
Representative top score
Not independently confirmed

Benchmarks to read alongside this one

METR Time Horizon: frequently asked questions

What is the METR time horizon?
It is a capability measure expressed in human time: the length of task, as timed with domain-expert humans, that a model completes with a 50% success rate. METR built it from RE-Bench, HCAST and 66 shorter tasks, publishing the method in March 2025.
How fast is the AI time horizon growing?
The original March 2025 paper reported a doubling roughly every seven months since 2019. The 2026 Time Horizon 1.1 update, which added longer tasks and removed flawed ones, put recent progress considerably faster than that original trend.
What does a 50% time horizon actually mean?
That the model succeeds on about half of the tasks of that human duration. It is a threshold, not a capability guarantee: a four-hour horizon means failing roughly half of four-hour tasks, which is why the number should not be read as "the AI can now do four hours of work."

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory