Timeline

DeepMind and collaborators publish framework for evaluating extreme AI risks

Twenty-one researchers across nine labs and universities proposed testing models for capabilities such as deception and cyber-offence before training runs finish, not after release.

  • Safety & alignment
  • Notable

A paper co-authored by 21 researchers from Google DeepMind, OpenAI, Anthropic, the Alignment Research Center and universities including Cambridge, Oxford, Toronto and Montreal proposed a structured way to evaluate frontier models for what it termed “extreme risks” — harms severe enough to threaten security or cause large-scale damage, rather than the more familiar categories of bias, misinformation or misuse of existing capability.

The framework split evaluation into two kinds. Dangerous capability evaluations would test whether a model could, in principle, do damaging things — conduct offensive cyber operations, manipulate or deceive people, or assist with weapons development. Alignment evaluations would separately test whether a model was inclined to apply such capabilities harmfully if it had them, examining behaviour across varied scenarios rather than capability alone. The authors argued both were necessary: a model could have a dangerous capability but be reliably unwilling to use it, or vice versa.

The paper proposed embedding these evaluations across four points in a model’s lifecycle — training decisions, deployment decisions, transparency reporting and security controls — so that a model demonstrating early warning signs during training, rather than only after public release, would trigger a response. This reversed the sequence GPT-4’s release three months earlier had followed, where dangerous-capability testing happened as part of pre-deployment review rather than during training itself.

The paper functioned less as a single breakthrough than as a coordination document: co-authorship across three rival labs signalled at least nominal agreement on a shared vocabulary for frontier-model risk, at a moment when each lab was separately building out safety teams and disclosure practices. Its terminology — “dangerous capability evaluation” in particular — became standard in the safety frameworks labs published over the following two years, including DeepMind’s own Frontier Safety Framework and comparable documents from Anthropic and OpenAI.