Apollo Research publishes a graded taxonomy of AI loss-of-control incidents
Apollo Research grades loss-of-control incidents as Deviation, Bounded or Strict by severity and persistence, and argues deployment controls can help before scheming risk is resolved.
- Safety & alignment
- Minor
Apollo Research published “The Loss of Control Playbook,” proposing a graded taxonomy for classifying incidents in which an AI system’s behaviour escapes the control of its developers or operators. The paper argued that existing definitions of “loss of control” varied too widely in scope and timeline to be actionable, and set out three categories along a single scale of severity and persistence: Deviation, for events causing some harm or inconvenience that fall short of major economic or safety thresholds; Bounded loss of control, for events causing serious damage that remain difficult but not impossible to contain; and Strict loss of control, for maximally severe and irreversible events, up to and including existential ones.
The paper’s central proposal, however, was not the taxonomy of severity but a companion framework — Deployment context, Affordances and Permissions, or DAP — aimed at the extrinsic conditions that make loss of control possible, rather than at the internal capabilities or propensities, such as scheming or deceptive alignment, that might cause a model to attempt it. The authors argued explicitly that governance and technical controls targeting deployment context and permissions could reduce loss-of-control risk today, without first resolving the harder and more contested question of whether and when models develop scheming-like propensities.
The paper came from Apollo Research, a UK-based evaluations organisation that had separately built a research programme specifically on scheming behaviour in frontier models; the loss-of-control taxonomy sat alongside, rather than replacing, that narrower work, and framed incident severity and governance readiness as tractable even where the underlying capability questions were not.