OpenAI releases GPT-5.4
OpenAI's first general-purpose model with built-in computer-use, reported scoring 75% on OSWorld-Verified against 47.3% for GPT-5.2 and roughly 72% for human testers.
- Models & capabilities
- Notable
OpenAI released GPT-5.4, shipped initially as two reasoning-focused variants, GPT-5.4 Thinking and GPT-5.4 Pro, two days after the non-reasoning GPT-5.3 Instant became ChatGPT’s new default quick-response model. Smaller GPT-5.4 mini and nano variants followed on 17 March, with mini made available to free-tier users and nano offered only through the API.
OpenAI presented GPT-5.4 as its first general-purpose model with computer-use capability built in natively, alongside a one-million-token context window and a tool-search mechanism the company said cut token costs by 47% in tool-heavy workflows. On OSWorld-Verified, a benchmark measuring an AI system’s ability to operate a desktop computer environment, OpenAI reported the model scoring 75%, against 47.3% for GPT-5.2 and an average human score of 72.4% — one of the few instances of a released model being cited as beating typical human performance on a real-world computer-operation task rather than a narrow academic one. The company also claimed a 33% reduction in factual errors compared with GPT-5.2. The accompanying system card described GPT-5.4 Thinking as the first general-purpose OpenAI model shipped with mitigations for high cybersecurity capability, building on measures introduced for GPT-5.3 Codex; some outside commentary noted the card gave less detailed treatment to safety evaluation of the model’s new computer-use capability itself than to its cyber capabilities.
Early coverage was mixed on specifics beneath the headline claims: reviewers praised lower hallucination rates and stronger research capability, but pricing drew criticism, since the mini and nano API variants launched roughly four times more expensive than their GPT-5 equivalents. As with other point releases in the series, the capability and cost-reduction figures came from OpenAI’s own evaluations rather than independent benchmarking.