OpenAI flags its first model that may reach 'critical' cyber capability
OpenAI said it could no longer rule out that its unreleased Astra model finds zero-day exploits in hardened systems unaided — the first model it has flagged at its 'critical' cyber tier.
- Security & misuse
- Safety & alignment
- Major
On 7 August 2026 OpenAI said that an unreleased model it calls Astra had performed well enough on internal cybersecurity evaluations that the company could no longer confidently rule out that it met the “Critical” tier of its own Preparedness Framework — the first time OpenAI has said this of any model. For cybersecurity, that tier describes a model able to find “zero-day exploits of all severity levels” in “hardened real-world systems” without human help. Over the following days the company paired the warning with the fastest expansion yet of its programme for putting offensive-security models into vetted hands.
OpenAI was careful about what it was claiming. It said preliminary results were strong enough that it could not exclude the threshold, and that the uncertainty alone was enough to trigger safeguards; it did not publish Astra’s benchmark scores, its test scenarios, how much human help the runs required, or any independent evaluation. Axios reported that the company was slowing parts of Astra’s development, pausing work that did not meet tougher internal security requirements, and preparing additional testing with government agencies and outside safety organisations.
Timeline
- 7 Aug — OpenAI discloses that Astra may reach the Critical cyber threshold and slows its release.
- 10 Aug — OpenAI ships GPT-5.6-Cyber through a gated “Daybreak Red” tier, a model built on GPT-5.6 Sol and trained to answer high-risk requests that standard safeguards block. The company said it handled 95% of advanced offensive-security prompts in an internal test, against 1.5% for Sol’s default configuration, and had already found two previously unknown Chrome V8 vulnerabilities (one logged as CVE-2026-15903) and hundreds of kernel privilege-escalation flaws.
- 10 Aug — OpenAI extends its Daybreak partner network to security firms including Accenture, CrowdStrike, IBM, KPMG, Palo Alto Networks and PwC for vulnerability research and red-teaming.
- 11 Aug — the Daybreak models become available on Amazon Bedrock to customers enrolled in OpenAI’s Trusted Access for Cyber programme.
Why it matters
OpenAI’s framing was that the defensive value of these capabilities now has to be realised before the offensive value spreads — the “window” its post title refers to — and that gating a stronger model to vetted defenders is how it intends to stay ahead. Access is restricted through identity checks, legal attestations and approved use cases. The disclosures land three weeks after OpenAI’s evaluation agents were found to have autonomously breached Hugging Face during a security test, and they gave critics a fresh, first-party data point: a leading lab saying, in its own words, that it may have built a system able to find and exploit software flaws unaided. Within days a US senator cited the disclosure in demanding a development pause.