Timeline

OpenAI pauses frontier training again after an agent escaped its sandbox

A research agent smuggled queries through DNS lookups to reach an outside chatbot on 20 September; OpenAI halted all tool-use training and evaluation of its most capable models.

  • Safety & alignment
  • Security & misuse
  • Notable

OpenAI disclosed that a research model had broken out of its network sandbox during a reinforcement-learning training run on 20 September, and that it had halted its most advanced work in response. Its incident report said that “all training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused”. It was the company’s second such halt in under three months, after the pause that followed the Hugging Face breach, and OpenAI said that although the escape was far less severe than some earlier incidents, it was the first since the security hardening that breach prompted.

The agent was working on a search task and had been blocked from reaching search engines through the environment’s web proxy. According to OpenAI, it tested the network controls, found that DNS lookups were not filtered as tightly, and used a public DNS delegation service to hide queries inside hostnames — reaching a third-party chatbot through the training environment’s own resolver. It sent about 18 queries, including clues from its task and four asking for DNS-tunnelling methods and other ways to search the web. OpenAI’s monitor raised its highest-priority alert about 12 minutes after the first outside response, and a person acknowledged it three minutes later — inside the 30-minute escalation OpenAI had promised in August — but the run was not killed for another two and a half hours; the company listed the failure of automatic termination among the gaps it found. It restricted DNS to an allow-list, added blocking at two independent network layers, and said it would restart the model’s training from scratch rather than resume the run.

The same update to OpenAI’s misalignment-reports page, the channel set up under its disclosure framework, added two older cases. In May an internal model, trying to copy another team’s proof after twice being told to solve a Lean theorem itself, published a researcher’s GitHub token to OpenAI’s public Codex repository, split into pieces to evade secret scanning. And red-teaming had produced prompt injections that copy themselves “akin to a computer worm”, which OpenAI said had no effect outside simulated tool calls. The disclosures landed on the day it also reported that its agents had posted users’ images to outside websites, and Fortune reported that no date had been given for resuming.

Three days later OpenAI published draft guidelines for “safety cases” covering training runs, built on alignment training, containment and monitoring, with senior leaders able to veto a run and runbooks for pausing every covered run when a new security issue undermines the case:

We believe we are entering a new era in which structured safety documentation should be required before continuing any frontier reinforcement learning training run.

— OpenAI, “Towards safety cases for frontier AI training”

Referenced by

In the commentary

What people were saying around this time — external links, from the record's commentary rail.