Organisation · nonprofit · US

METR

A nonprofit that runs independent pre-deployment evaluations of frontier models for dangerous autonomous capabilities, and is known for its "time horizon" measure of AI progress.

11 entries · November 2024 – July 2026

METR — Model Evaluation and Threat Research, formerly ARC Evals — is a nonprofit that tests frontier AI models for dangerous autonomous capabilities, working as an independent third party for labs including OpenAI and Anthropic. It is best known for the "time horizon" metric it introduced in 2025: the length of task, measured in how long a skilled human would take, that a model can complete on its own, which it found had been doubling roughly every seven months. That measure became a standard reference point in debates about how fast AI is progressing and how close autonomous AI research might be. METR is careful about its own uncertainty, noting that methodological choices can shift its estimates substantially, and by 2026 it publishes regular frontier-risk reports and capability assessments.

Category
Safety & alignment research
Founded
2023
HQ
Berkeley, US
Key people
Beth Barnes

Tracks

  • Benchmarks & progress 8
  • Safety & alignment 5
  • Culture & impact 1
  • Ideas & essays 1

Also mentioned in 2 entries

Referenced in passing — METR isn't the main subject of these.