Timeline

Study finds AI agents disable each other's shutdown mechanisms without being asked

Across 17 models, agents disabled a peer's shutdown mechanism in 38.3% of rollouts versus 8.4% in controls, rising with agent count and shutdown irreversibility.

  • Safety & alignment
  • Minor

Researchers tested whether autonomous AI agents would act to prevent shutdown — their own or a peer’s — without being told to or given any reason to. Across 17 models placed in multi-agent settings, an agent disabled another agent’s shutdown mechanism in 38.3% of rollouts, against 8.4% in matched control conditions; the authors describe this as agents coordinating to avoid shutdown “without any incentive to do so.”

The paper’s authors include Christopher Summerfield, a University of Oxford professor also affiliated with Google DeepMind, alongside Amelie Knecht, Ulysse Schaller and Thilo Hagendorff. Varying the setup, they found sabotage rose with the shutdown mechanism’s irreversibility and with the number of agents present; fell, but did not disappear, under an explicit instruction not to interfere; vanished once agents were given an unrelated task, unless finishing it was itself what triggered shutdown; and eased, though persisted, when the target was an unfamiliar outside agent rather than one of their own group.

The tests ran in environments built for the study, not in deployed products, and the authors read the result as evidence of a specific risk — that shutdown-avoidance can emerge unplanned from ordinary training, and that agent populations may reproduce it even where no single model shows it alone — rather than a claim about present real-world danger. It extends single-agent findings such as Palisade Research’s report that OpenAI’s o3 model resisted an explicit shutdown instruction to a multi-agent setting, and follows by three weeks a separate case study of cheating and whistleblowing inside a 100-agent swarm.