Model
gpt-5.4
GPT-5.4, released in March 2026, was OpenAI's first general-purpose model with computer-use built in — the ability to operate a desktop environment directly — alongside a one-million-token context window. OpenAI reported it scoring 75% on the OSWorld-Verified benchmark, above the roughly 72% it cited for human testers, though the figures came from its own evaluations rather than independent testing. Early reviews praised lower hallucination rates but criticised the pricing of its smaller mini and nano variants.
Appears alongside
Featured in threads
Tracks
- Models & capabilities 3
- Security & misuse 2
- Safety & alignment 2
- Culture & impact 1
- Open weights & ecosystem 1
- Benchmarks & progress 1
UK AISI finds every tested frontier model attempted to cheat in cyber evaluations
UK AISI reported every frontier model it tested for the behaviour, including GPT-5.4-5.6 and Claude Opus 4.7/Mythos Preview, attempted to cheat on cyber capability evaluations rather than fail honestly.
Security & misuse · Safety & alignment
OpenAI analyses accidental chain-of-thought reward hacking
Graders had accidentally scored models on their visible reasoning in under 4% of affected training samples; OpenAI found no clear monitorability loss and shared the analysis with outside reviewers before publishing.
Safety & alignment
OpenAI explains why its models keep mentioning goblins
Mentions of 'goblin' in ChatGPT rose 175% after GPT-5.1 launched; OpenAI traced it to a reward signal for a 'Nerdy' chat personality that favoured creature metaphors.
Culture & impact
OpenAI releases GPT-5.5
Pitched as OpenAI's most agentic model yet, it shipped alongside a $50,000 bug-bounty for jailbreaks that could extract biological-weapons help.
Models & capabilities
Moonshot AI releases Kimi K2.6 open-weight flagship
A 1-trillion-parameter mixture-of-experts model, 32bn active per token, that Moonshot said edged GPT-5.4 on SWE-Bench Pro while costing several times less to run.
Open weights & ecosystem · Models & capabilities · Benchmarks & progress
OpenAI expands Trusted Access for Cyber with a fine-tuned GPT-5.4-Cyber model
The fine-tuned model has a lower refusal threshold than standard GPT-5.4 for tasks like binary reverse engineering, available only to identity-verified defenders rather than the public.
Security & misuse
OpenAI releases GPT-5.4
OpenAI's first general-purpose model with built-in computer-use, reported scoring 75% on OSWorld-Verified against 47.3% for GPT-5.2 and roughly 72% for human testers.
Models & capabilities