Timeline

METR reports frontier models show dangerous capability before public deployment

METR argued that model theft, internal misuse and misaligned agents pose risks during training and internal deployment, before any public release.

  • Safety & alignment
  • Minor

METR, the nonprofit whose capability evaluations feed into pre-release testing at labs including OpenAI and Anthropic, argued in a blog post that the industry’s safety practice — built around evaluating a model just before it is made public — misses risks that arise earlier, while a model is still in training or in use only inside its own developer.

The post identified three risks pre-deployment testing does not reach: theft of model weights or algorithmic secrets by outside actors, catastrophic misuse by employees with internal access, and misaligned agents pursuing harmful goals autonomously during training runs or internal deployment, none of which an external evaluator working from a finished system card would see. METR drew a comparison to biosecurity and nuclear research, where security obligations apply “long before the technologies are actually deployed,” and argued frontier AI development lacked an equivalent regime.

Rather than call simply for longer pre-release testing windows, METR recommended earlier capability forecasting during training itself, stronger internal security and monitoring at labs, transparency and responsible-disclosure mechanisms, and stronger protections for employees who raised concerns.

The post did not point to a specific incident that had already occurred; it was a structural argument about where the industry’s evaluation regime, largely organised around system cards published at public launch, had a gap — one that became a reference point in later debate over what labs owed regulators and the public about what happened before release.