Timeline

METR examines how time horizon varies across domains

Applying its 50%-success task-length method to nine benchmarks, METR found doubling times of two to six months for reasoning tasks but around twenty months for Tesla's self-driving system.

  • Benchmarks & progress
  • Minor

METR extended the “time horizon” methodology from its earlier work — the length of task, in human-time terms, that a model can complete autonomously with 50% success probability — to nine existing benchmarks spanning software engineering (METR-HRS, SWE-bench, LiveCodeBench), scientific and mathematical reasoning (GPQA Diamond, MATH, Mock AIME), computer use (OSWorld, WebArena), and Tesla’s Full Self-Driving system, estimating time horizons from publicly available model performance data rather than running new experiments.

The results varied sharply by domain but agreed on one point: no domain showed growth slowing down. Intellectual tasks — maths, coding, scientific question-answering — showed horizons of 50 to 200-plus minutes doubling every two to six months. Agentic computer-use tasks showed horizons roughly 40 to 100 times shorter than the intellectual-task domains but growing at broadly similar rates. Tesla’s self-driving system stood out as the clear outlier, improving at only around 0.6 doublings a year — a doubling time closer to twenty months, far slower than any of the reasoning benchmarks.

METR flagged significant limitations: the nine benchmarks cover only a narrow slice of real economic activity, excluding domains such as management or nursing; the tasks themselves are generally easier and more artificial than their real-world equivalents; human baseline times often rely on estimates rather than measurement; and most of the evidence comes from domains where AI models already perform comparatively well, which could bias the overall picture toward faster apparent progress than a genuinely representative task sample would show.