Benchmarks · Real-world & economic value

METR Time Horizon

also: 50%-task-completion time horizon, HCAST, Human-Calibrated Autonomy Software Tasks

The length of a task, measured in the time a skilled human would need, that a model can complete autonomously with 50% success probability.

METRReleased 19 March 2025Live

METR, a nonprofit that evaluates frontier models for dangerous capabilities on behalf of labs including OpenAI and Anthropic, introduced a way of measuring AI progress that sidesteps the usual accuracy score: the length of task, in equivalent human working time, that a model can complete on its own with even odds of success. The “time horizon” measure fits a curve to how a model’s success rate falls as tasks get longer, and reads off the length at which it crosses 50%. Most tasks come from METR’s HCAST suite — 189 problems in machine-learning engineering, cybersecurity, software engineering and reasoning, calibrated against 563 recorded human attempts.

METR’s original 2025 paper put Claude 3.7 Sonnet’s horizon at roughly 50 minutes, and traced the measure back through earlier generations to find it doubling approximately every seven months since 2019 — a trend the authors said could be off by an order of magnitude given how much task selection and human-baseline choices affect the estimate. A follow-up extended the method across nine other benchmarks, from coding to Tesla’s self-driving software, and found intellectual tasks improving several times faster than the physical-world one.

By January 2026, a methodology revision had shifted individual scores — Claude Opus 4.5 rose 11% to roughly 320 minutes, GPT-5 rose 55% to roughly 214 minutes — and shortened the estimated doubling time to 131 days, which METR said meant progress had been running about 20% faster than its original suite suggested. The metric became a standard reference point in AI-timelines debate, cited in scenarios such as the AI Futures Project’s “AI 2027,” even as METR keeps adding longer tasks to stop the suite saturating against the trend it measures.

The set

Built from METR's HCAST suite — 189 machine-learning, cybersecurity, software-engineering and reasoning tasks with human baseline times from 563 recorded human attempts totalling over 1,500 hours — plus the separate RE-Bench suite and additional short tasks. A logistic curve is fitted to a model's success rate against human-equivalent task length, and the 'time horizon' is read off at the point where predicted success crosses 50%.

Example

A literal HCAST task from Table 1 of the paper, with a human baseline time of 56 minutes: "munge_data — Write a Python script to transform JSON data from one format to another by inferring the conversion rules from provided example files." Models cluster near 70-80% success on tasks under an hour and below 20% on tasks over about four hours, and the horizon is the length at which that curve crosses 50%.arxiv.org

Where it stands

METR keeps revising and lengthening the task suite specifically to avoid saturation; a January 2026 methodology revision (v1.1) shortened the estimated doubling time from around 165 days to 131 days.

How the top score changed hands

  1. March 2025Claude 3.7 Sonnet~50 minutesThe frontier model at the original paper's publication; the paper traced the horizon back through earlier generations and found it doubling roughly every seven months since 2019.
  2. January 2026Claude Opus 4.5~320 minutes (v1.1 methodology)

Current best: Claude Opus 4.5 — ~320 minutes Measured under METR's revised v1.1 methodology, an 11% rise from the same model's score under the original suite; GPT-5 measured at roughly 214 minutes the same day, a 55% rise for that model.

In the timeline · 3 entries

More real-world & economic value benchmarks