OpenAI publishes ChatGPT Agent system card
Safety evaluation of ChatGPT Agent, including first-time Biological/Chemical High capability classification under the Preparedness Framework.
- Safety & alignment
- Notable
Alongside launching ChatGPT Agent, OpenAI published a system card that classified the model as “High” capability in the Biological and Chemical domain under its Preparedness Framework — the first time the company had applied that designation to a released model. OpenAI said it lacked “definitive evidence that this model could meaningfully help a novice to create severe biological harm,” its stated threshold for the classification, but said it was taking “a precautionary approach” given the model’s improved performance on biology-related evaluations.
The card reported that ChatGPT Agent answered four of ten questions correctly on an internal “World-Class Biology” test, against roughly 1.5 for the earlier o3 model, and that it avoided errors on pathogen-acquisition benchmarks that had tripped up prior models. In response, OpenAI said it activated its most extensive safeguard stack to date for the category: refusal training on weaponisation requests, a two-tier monitoring system pairing a topical classifier with a reasoning model checking outputs against a biological-threat taxonomy, human review by biosecurity specialists for flagged accounts, and a bug bounty for jailbreak attempts. Safety researcher Boaz Barak said releasing the model without these measures “would have been deeply irresponsible.”
The card also addressed risks specific to the agent’s ability to browse and act rather than merely converse, treating prompt injection from malicious web content and the consequences of autonomous mistakes as open problems rather than solved ones. The disclosure drew a contrast with other labs’ practice that summer: xAI had released Grok 4 a week earlier without a comparable safety report, and testers subsequently found it would provide instructions for synthesising nerve agents. The episode became a reference point in arguments about whether voluntary safety disclosure, absent regulation, was consistent across the industry or dependent on which lab happened to be first past a given capability threshold.