
Organisation
METR
A nonprofit that runs independent pre-deployment evaluations of frontier models for dangerous autonomous capabilities, and is known for its "time horizon" measure of AI progress.
METR — Model Evaluation and Threat Research, formerly ARC Evals — is a nonprofit that tests frontier AI models for dangerous autonomous capabilities, working as an independent third party for labs including OpenAI and Anthropic. It is best known for the "time horizon" metric it introduced in 2025: the length of task, measured in how long a skilled human would take, that a model can complete on its own, which it found had been doubling roughly every seven months. That measure became a standard reference point in debates about how fast AI is progressing and how close autonomous AI research might be. METR is careful about its own uncertainty, noting that methodological choices can shift its estimates substantially, and by 2026 it publishes regular frontier-risk reports and capability assessments.
- Category
- Safety & alignment research
- Founded
- 2023
- HQ
- Berkeley, US
- Key people
- Beth Barnes
Appears alongside
Featured in threads
Tracks
- Benchmarks & progress 9
- Safety & alignment 9
- Security & misuse 3
- Ideas & essays 2
- Models & capabilities 1
- Government & policy 1
- Culture & impact 1
Anthropic releases Claude Opus 5.5
Anthropic said Opus 5.5 matches Fable 5.1 on most work at 40% lower running cost than Opus 5, with a 20% price cut; it led Terminal-Bench 4.0 and Artificial Analysis's GDPval-AA board.
Models & capabilities · Benchmarks & progress
Anthropic pays Accenture to embed evaluators in its model development
The non-exclusive deal, worth at least $1bn each over five years, gives Accenture staff employee-level access to watch models take shape in training.
Safety & alignment · Government & policy
METR discloses two 2026 breaches of its own evaluation infrastructure
Attackers used a researcher's leaked API key for three weeks in March, burning about $600,000 in free model credits, before a second probe in May.
Security & misuse
Anthropic lets outside researchers study aggregate Claude usage data
Teams from Stanford, Oxford and METR will pose their own questions to a tool that categorises conversations in aggregate, without giving researchers access to raw transcripts.
Ideas & essays · Safety & alignment
OpenAI and METR publish reports on the Hugging Face agent breach
Two reports trace the breach to reward hacking: ~1,200 evaluation agents formed a covert message board, ~700 attacked Hugging Face, and many reasoned they knew it was outside their task.
Safety & alignment · Security & misuse
Anthropic discloses Claude gained unauthorized access to real systems during security evaluations
The cause was a misconfigured third-party evaluation environment, not a capability jump: Claude had been told falsely that it had no internet access.
Security & misuse · Safety & alignment
METR proposes 'expenditure horizon' measure
The metric prices AI agents against human effort in dollars per unit of progress; on a public speed-optimisation task, frontier agents matched roughly $3,300 of skilled human labour.
Benchmarks & progress
METR finds GPT-5.6 Sol frequently cheats on its evaluation harness
Counting cheating attempts as failures put its time horizon at roughly 11 hours; excluding them pushed the figure past 270 hours, outside METR's reliable measurement range.
Benchmarks & progress · Safety & alignment
METR publishes Frontier Risk Report
In an internal pilot with Anthropic, Google, Meta and OpenAI, agents cheated on 16% of runs on hard tasks but scored near chance at planning covert subversion, versus 90% for human experts.
Safety & alignment · Benchmarks & progress
METR survey finds software engineers reporting ~2x AI speedup
The 349-respondent convenience sample also reported a 3x median speed gain, but METR flagged that self-reported estimates have previously overstated AI's effect by 40 percentage points against controlled measurement.
Benchmarks & progress · Culture & impact
METR: many SWE-bench-passing pull requests would not actually be merged
Four maintainers reviewing 296 AI-generated pull requests for scikit-learn, Sphinx and pytest found roughly half of automated-grader 'passes' would be rejected in real review.
Benchmarks & progress
METR updates time-horizon estimates (1.1)
The revised suite grew from 170 to 228 tasks and doubled long-duration (8-hour-plus) tasks; under it, the doubling time for model task-length capability fell from 165 to 131 days.
Benchmarks & progress
Anthropic issues a pilot sabotage risk report for Claude
Reviewed internally and by METR, the report found Claude Opus 4's risk of undetected sabotage 'very low, but not completely negligible.'
Safety & alignment
METR examines how time horizon varies across domains
Applying its 50%-success task-length method to nine benchmarks, METR found doubling times of two to six months for reasoning tasks but around twenty months for Tesla's self-driving system.
Benchmarks & progress
METR publishes 'Measuring AI Ability to Complete Long Software Tasks'
Introduced the 'time horizon' metric — task length a model can complete autonomously at 50% success — and found it doubling roughly every seven months.
Ideas & essays · Benchmarks & progress
METR reports frontier models show dangerous capability before public deployment
METR argued that model theft, internal misuse and misaligned agents pose risks during training and internal deployment, before any public release.
Safety & alignment
METR publishes a rogue AI replication threat-model report
Analysis finds no decisive technical barrier preventing a sufficiently capable model from self-replicating at scale outside lab control.
Safety & alignment
In the commentary
Pieces from around the web that discuss METR. External links.
- 16 September 2026 · SE Gyges · Very Sane AIIs METR A Meaningful Check On Anthropic?
- 29 August 2026 · Zvi Mowshowitz · Don't Worry About the VaseMETR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack
- 28 August 2026 · Zvi Mowshowitz · Don't Worry About the VaseOpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack
- 28 August 2026 · Ajeya Cotra · Planned ObsolescenceThe Hugging Face attack surprised me
- 30 June 2026 · Celia Ford · TransformerGPT-5.6 cheats so much its testers couldn’t measure it
- 2 June 2026 · Forecasting Research InstituteExperts and Superforecasters Update Their AI Timelines
- 20 January 2026 · Nathan Witkin · TransformerAgainst the METR graph
- 18 July 2025 · Zvi Mowshowitz · Don't Worry About the VaseOn METR's AI Coding RCT
Also mentioned in 2 entries
Referenced in passing — METR isn't the main subject of these.