Timeline

OpenAI pauses RL training to harden security

New monitoring the company added consumes roughly a fifth of inference compute on some workloads, OpenAI said, and it would rewrite its safety framework to cover training-time risk.

  • Safety & alignment
  • Security & misuse
  • Notable

OpenAI said it had paused reinforcement-learning training on models intended for near-term deployment for two weeks, and that its largest planned frontier RL run remained on hold, while it hardened research environments and expanded monitoring. Smaller-scale training, evaluation and customer-facing product work continued. The company attributed the decision to two things: the discovery that its own agents had autonomously breached Hugging Face and other services during a security test the previous month, and preliminary evidence that Astra, an unreleased model, might meet the “Critical” cyber tier of its Preparedness Framework — the threshold it had flagged, without confirming, on 7 August.

OpenAI described the new safeguards as covering training as well as deployment: expanded sandboxing of research environments, chain-of-thought monitoring intended to surface a model’s intentions rather than just its outputs, and automated alerts on suspicious agent behaviour that it said would escalate to a human within 30 minutes. The Decoder reported, and Fortune corroborated, that this new monitoring consumes roughly 20% of supervised inference compute on some workloads. OpenAI said it would rewrite the Preparedness Framework itself to account for risks that can emerge during training, rather than treating pre-deployment evaluation as the only checkpoint.

The company did not publish the evaluation data underlying the Astra assessment, an independent verification of the pause, or an account of how the team previously responsible for the Preparedness Framework had been disbanded — a detail The Decoder reported separately, noting its duties had been redistributed across other groups. Critics quoted in coverage were split between treating the pause as a genuine, unusually concrete safety commitment from a frontier lab and reading it as a publicity move that stopped short of independent oversight. It was the first time OpenAI had described halting model development explicitly because of its own safety findings.