Timeline

OpenAI publishes its Preparedness Framework

The beta framework scored models low to critical on four risk categories, barring deployment above 'high' and barring further development above 'critical,' with the board holding final oversight.

  • Safety & alignment
  • Notable

OpenAI published, in beta, its Preparedness Framework, a formal process for evaluating and responding to severe risks from its own frontier models — three months after Anthropic’s Responsible Scaling Policy established the template of a self-governing, capability-triggered safety commitment, and about a month after the board briefly removed and reinstated Sam Altman in a crisis that had turned public attention onto how the company governed itself.

The framework tracked four categories of catastrophic risk — cybersecurity, chemical/biological/radiological/nuclear (CBRN) capability, persuasion, and model autonomy — and scored each model’s capability in every category as low, medium, high or critical, based on evaluations OpenAI’s Preparedness team ran before deployment. Under the framework’s rules, a model scoring “high” in any category could not be deployed until mitigations brought the residual risk down; a model scoring “critical” halted further development entirely, not just release, until the risk was addressed. An internal Safety Advisory Group reviewed capability reports and issued recommendations to leadership, with OpenAI’s board retaining oversight authority, including the ability to reverse a leadership deployment decision.

Like Anthropic’s policy, the framework was voluntary and self-administered: no external body was empowered to enforce its thresholds, and OpenAI itself decided when a mitigation was sufficient. Its publication nonetheless marked the framework’s translation from research proposal into a stated operational commitment inside one of the two leading labs, and it joined the RSP as a reference point other companies and, later, policymakers cited when describing what industry self-regulation on frontier risk could look like. OpenAI substantially revised the document as Version 2 in April 2025, collapsing the four-tier scale into two operational thresholds and adding a clause allowing safeguards to be relaxed if a competitor shipped a comparably risky model without them.