Timeline

DeepMind publishes the Frontier Safety Framework

A set of internal capability thresholds across autonomy, cybersecurity, biosecurity and ML R&D, joining similar voluntary policies already published by Anthropic and OpenAI.

  • Safety & alignment
  • Major

Google DeepMind published the first version of its Frontier Safety Framework, a set of protocols for identifying when a model’s capabilities might pose severe risk before that risk materialises. The framework’s central device is the Critical Capability Level (CCL): a threshold describing the minimum capability a model would need to meaningfully contribute to a specific category of harm. DeepMind said it would run “early warning evaluations” designed to detect when a model was approaching a CCL well before it crossed one, giving time to apply mitigations.

The initial version specified four risk domains: autonomous replication and adaptation, cybersecurity, biosecurity, and machine learning research and development capable of accelerating a model’s own further development. The document set out what DeepMind would do on breach — restricting deployment, tightening security around model weights — without publishing the specific capability thresholds themselves, and it committed to revisiting the framework as the science of evaluation matured.

The publication placed DeepMind alongside Anthropic’s Responsible Scaling Policy (September 2023) and OpenAI’s Preparedness Framework (December 2023) as the third of the major labs to commit publicly to capability-triggered safety measures rather than blanket policies applied regardless of what a model could do. All three frameworks were voluntary and self-administered — none created an external body with power to block a release — which made them, in effect, a bet that the labs writing the rules would also enforce them against their own commercial incentives. The framework arrived the same week as Jan Leike’s resignation from OpenAI and his public complaint that safety work there had “taken a backseat to shiny products,” a contrast DeepMind’s announcement did not draw but commentators did.

Referenced by