METR updates time-horizon estimates (1.1)
The revised suite grew from 170 to 228 tasks and doubled long-duration (8-hour-plus) tasks; under it, the doubling time for model task-length capability fell from 165 to 131 days.
- Benchmarks & progress
- Minor
METR published version 1.1 of its time-horizon benchmark, which measures the length of task, in estimated human working time, that a model can complete with 50% reliability. The update expanded the task suite from 170 to 228 tasks — adding 73, dropping 15 it judged unreliable, and revising 53 others — and doubled the number of long-duration tasks requiring eight or more hours of human effort, from 14 to 31, narrowing the confidence intervals on frontier models’ scores. METR also migrated its evaluation infrastructure from its in-house Vivaria system to Inspect, an open-source framework built by the UK AI Security Institute.
The new methodology shifted individual model scores: Claude Opus 4.5’s estimated time horizon rose 11% to roughly 320 minutes and GPT-5’s rose 55% to roughly 214 minutes, while several older GPT-4-era models’ scores fell by 35–57%, changes METR attributed to the revised task mix rather than any change in the underlying models. Recalculated under the new suite, the time it takes for models’ time horizons to double fell from roughly 165 days (measured from 2023 onward under the original methodology) to 131 days — implying, METR said, that progress had been running about 20% faster than the original benchmark suggested. Measured only from 2024 onward, the doubling time was shorter still, at roughly 89 days.
METR cautioned that its confidence intervals remained wide and that it was continuing to add longer tasks to avoid the suite saturating as models improve — an explicit acknowledgement that any fixed benchmark of this kind has a limited shelf life against models that keep gaining on it.