Timeline

DeepMind expands the Frontier Safety Framework to cover manipulation and shutdown resistance

Version 3.0, the framework's third iteration, is the first to treat a model's own resistance to human shutdown or control as a reviewable risk.

  • Safety & alignment
  • Notable

Google DeepMind published version 3.0 of its Frontier Safety Framework, the third revision of the Critical Capability Level system it introduced in May 2024 to flag when a model’s abilities cross thresholds serious enough to warrant restricted deployment.

The update added a new Critical Capability Level for harmful manipulation: models judged capable of systematically and substantially changing beliefs and behaviour at scale in high-stakes settings would now trigger the framework’s mitigations, joining the existing domains of autonomous replication, cybersecurity, biosecurity and ML research and development. Authored by Four Flynn, Helen King and Anca Dragan, the update also expanded governance around misalignment — cases where a model might interfere with an operator’s ability to direct, modify or shut it down — requiring a formal safety case review before external launch and extending oversight to large-scale internal deployments as well, not only public releases.

As with earlier versions, the framework remained voluntary and self-administered: DeepMind set its own thresholds, ran its own evaluations, and decided for itself when a mitigation had been satisfied. No external body was given power to compel a delay or a launch decision. The addition of shutdown resistance as a named risk category was nonetheless notable for treating a model’s disposition toward its own operators — rather than only its capacity to help a third party cause harm — as something worth formally testing for ahead of release, a question that had mostly been debated in research papers on scheming and deceptive alignment rather than written into a lab’s operational policy.

DeepMind revised the framework again the following April, adding a lower-severity “Tracked Capability Level” tier intended to catch risks earlier than a full Critical Capability Level threshold would.