OpenAI publishes GPT-5 system card
OpenAI classified the reasoning variant as High capability for biological and chemical risk under its Preparedness Framework, its first model to reach that tier in the category.
- Safety & alignment
- Notable
OpenAI published a system card alongside the release of GPT-5, covering safety evaluation of the model’s variants, including the reasoning-focused gpt-5-thinking. The document said OpenAI would treat gpt-5-thinking as High capability in the biological and chemical domain under its Preparedness Framework — the tier below Critical, at which the framework requires stronger deployment safeguards. It was the first OpenAI model rated High in that category, activating restrictions such as tighter content filtering and expanded monitoring for prompts touching on pathogen or toxin synthesis. The model was not rated High for cybersecurity capability; OpenAI reached that threshold for a model only later, with GPT-5.5 and its Codex variants.
The card described a red-teaming programme OpenAI said comprised more than 5,000 hours of work from external testers, covering violent attack planning, jailbreaks, prompt injection and bioweapon-related queries, alongside OpenAI’s standard evaluation suite for persuasion, deception and autonomous replication. It also included the results underpinning the model’s safe-completions training, the output-focused approach to handling ambiguous dual-use prompts that replaced binary refusal.
As with previous system cards, the evaluations were run and interpreted by OpenAI itself, with third-party testers contracted rather than independent of the company. The document set the template other labs’ summer 2025 releases were compared against, and the biological-risk classification became a reference point in later debate over whether frontier labs’ self-assessments under the Preparedness Framework were keeping pace with actual capability gains.