Benchmarks · Real-world & economic value
METR Time Horizon
also: 50%-task-completion time horizon, HCAST, Human-Calibrated Autonomy Software Tasks
The length of a task, measured in the time a skilled human would need, that a model can complete autonomously with 50% success probability.
METRReleased 19 March 2025Live
METR, a nonprofit that evaluates frontier models for dangerous capabilities on behalf of labs including OpenAI and Anthropic, introduced a way of measuring AI progress that sidesteps the usual accuracy score: the length of task, in equivalent human working time, that a model can complete on its own with even odds of success. The “time horizon” measure fits a curve to how a model’s success rate falls as tasks get longer, and reads off the length at which it crosses 50%. Most tasks come from METR’s HCAST suite — 189 problems in machine-learning engineering, cybersecurity, software engineering and reasoning, calibrated against 563 recorded human attempts.
METR’s original 2025 paper put Claude 3.7 Sonnet’s horizon at roughly 50 minutes, and traced the measure back through earlier generations to find it doubling approximately every seven months since 2019 — a trend the authors said could be off by an order of magnitude given how much task selection and human-baseline choices affect the estimate. A follow-up extended the method across nine other benchmarks, from coding to Tesla’s self-driving software, and found intellectual tasks improving several times faster than the physical-world one.
By January 2026, a methodology revision had shifted individual scores — Claude Opus 4.5 rose 11% to roughly 320 minutes, GPT-5 rose 55% to roughly 214 minutes — and shortened the estimated doubling time to 131 days, which METR said meant progress had been running about 20% faster than its original suite suggested. The metric became a standard reference point in AI-timelines debate, cited in scenarios such as the AI Futures Project’s “AI 2027,” even as METR keeps adding longer tasks to stop the suite saturating against the trend it measures.
The set
Built from METR's HCAST suite — 189 machine-learning, cybersecurity, software-engineering and reasoning tasks with human baseline times from 563 recorded human attempts totalling over 1,500 hours — plus the separate RE-Bench suite and additional short tasks. A logistic curve is fitted to a model's success rate against human-equivalent task length, and the 'time horizon' is read off at the point where predicted success crosses 50%.
Example
A literal HCAST task from Table 1 of the paper, with a human baseline time of 56 minutes: "munge_data — Write a Python script to transform JSON data from one format to another by inferring the conversion rules from provided example files." Models cluster near 70-80% success on tasks under an hour and below 20% on tasks over about four hours, and the horizon is the length at which that curve crosses 50%.arxiv.org
Where it stands
METR keeps revising and lengthening the task suite specifically to avoid saturation; a January 2026 methodology revision (v1.1) shortened the estimated doubling time from around 165 days to 131 days.
How the top score changed hands
- March 2025Claude 3.7 Sonnet~50 minutesThe frontier model at the original paper's publication; the paper traced the horizon back through earlier generations and found it doubling roughly every seven months since 2019.
- January 2026Claude Opus 4.5~320 minutes (v1.1 methodology)
Current best: Claude Opus 4.5 — ~320 minutes Measured under METR's revised v1.1 methodology, an 11% rise from the same model's score under the original suite; GPT-5 measured at roughly 214 minutes the same day, a 55% rise for that model.
In the timeline · 3 entries
METR updates time-horizon estimates (1.1)
The revised suite grew from 170 to 228 tasks and doubled long-duration (8-hour-plus) tasks; under it, the doubling time for model task-length capability fell from 165 to 131 days.
Benchmarks & progress
METR examines how time horizon varies across domains
Applying its 50%-success task-length method to nine benchmarks, METR found doubling times of two to six months for reasoning tasks but around twenty months for Tesla's self-driving system.
Benchmarks & progress
METR publishes 'Measuring AI Ability to Complete Long Software Tasks'
Introduced the 'time horizon' metric — task length a model can complete autonomously at 50% success — and found it doubling roughly every seven months.
Ideas & essays · Benchmarks & progress