Timeline

METR discloses two 2026 breaches of its own evaluation infrastructure

Attackers used a researcher's leaked API key for three weeks in March, burning about $600,000 in free model credits, before a second probe in May.

  • Security & misuse
  • Minor

METR, the nonprofit that evaluates frontier models for dangerous capability, disclosed two security incidents against its own infrastructure earlier in 2026. In March, an attacker found an API key exposed on a researcher’s personal cloud instance — located by searching certificate-transparency logs for AI-related keywords — and exploited a fail-open authentication flaw to use it for roughly three weeks, consuming about $600,000 worth of model credits that METR said had been granted to it for free. In May, a separate, coordinated campaign probed METR’s public infrastructure through credential stuffing, token abuse and phishing, alongside an inadvertently exposed query path in its public transcript viewer.

METR said no sensitive information was accessed in either incident, though some model-output data was “inadvertently accessible in principle” in May, with no sign attackers found the exposure.

The disclosure is notable less for its scale than its target: one of the organisations tasked with independently assessing AI risk being attacked itself, at a time when METR had taken on a higher public profile through work including its joint investigation into the OpenAI–Hugging Face agent breach. It added evaluators to the list of infrastructure worth securing against the same automated, credential-driven probing AI-security researchers spent 2026 warning about.

In the commentary

What people were saying around this time — external links, from the record's commentary rail.