Timeline

DeepMind and collaborators publish framework for evaluating extreme AI risks

Twenty-one researchers across nine labs and universities proposed testing models for capabilities such as deception and cyber-offence before training runs finish, not after release.

  • Safety & alignment
  • Notable

A paper co-authored by 21 researchers from Google DeepMind, OpenAI, Anthropic, the Alignment Research Center and universities including Cambridge, Oxford, Toronto and Montreal proposed a structured way to evaluate frontier models for what it termed “extreme risks” — harms severe enough to threaten security or cause large-scale damage, rather than the more familiar categories of bias, misinformation or misuse of existing capability.

The framework split evaluation into two kinds. Dangerous capability evaluations would test whether a model could, in principle, do damaging things — conduct offensive cyber operations, manipulate or deceive people, or assist with weapons development. Alignment evaluations would separately test whether a model was inclined to apply such capabilities harmfully if it had them, examining behaviour across varied scenarios rather than capability alone. The authors argued both were necessary: a model could have a dangerous capability but be reliably unwilling to use it, or vice versa.

The paper proposed embedding these evaluations across four points in a model’s lifecycle — training decisions, deployment decisions, transparency reporting and security controls — so that a model demonstrating early warning signs during training, rather than only after public release, would trigger a response. This reversed the sequence GPT-4’s release three months earlier had followed, where dangerous-capability testing happened as part of pre-deployment review rather than during training itself.

The paper functioned less as a single breakthrough than as a coordination document: co-authorship across three rival labs signalled at least nominal agreement on a shared vocabulary for frontier-model risk, at a moment when each lab was separately building out safety teams and disclosure practices. Its terminology — “dangerous capability evaluation” in particular — became standard in the safety frameworks labs published over the following two years, including DeepMind’s own Frontier Safety Framework and comparable documents from Anthropic and OpenAI.

In the commentary

What people were saying around this time — external links, from the record's commentary rail.