Timeline

OpenAI flags its first model that may reach 'critical' cyber capability

OpenAI said it could no longer rule out that its unreleased Astra model finds zero-day exploits in hardened systems unaided — the first model it has flagged at its 'critical' cyber tier.

  • Security & misuse
  • Safety & alignment
  • Major

On 7 August 2026 OpenAI said that an unreleased model it calls Astra had performed well enough on internal cybersecurity evaluations that the company could no longer confidently rule out that it met the “Critical” tier of its own Preparedness Framework — the first time OpenAI has said this of any model. For cybersecurity, that tier describes a model able to find “zero-day exploits of all severity levels” in “hardened real-world systems” without human help. Over the following days the company paired the warning with the fastest expansion yet of its programme for putting offensive-security models into vetted hands.

OpenAI was careful about what it was claiming. It said preliminary results were strong enough that it could not exclude the threshold, and that the uncertainty alone was enough to trigger safeguards; it did not publish Astra’s benchmark scores, its test scenarios, how much human help the runs required, or any independent evaluation. Axios reported that the company was slowing parts of Astra’s development, pausing work that did not meet tougher internal security requirements, and preparing additional testing with government agencies and outside safety organisations.

Timeline

  • 7 Aug — OpenAI discloses that Astra may reach the Critical cyber threshold and slows its release.
  • 10 Aug — OpenAI ships GPT-5.6-Cyber through a gated “Daybreak Red” tier, a model built on GPT-5.6 Sol and trained to answer high-risk requests that standard safeguards block. The company said it handled 95% of advanced offensive-security prompts in an internal test, against 1.5% for Sol’s default configuration, and had already found two previously unknown Chrome V8 vulnerabilities (one logged as CVE-2026-15903) and hundreds of kernel privilege-escalation flaws.
  • 10 Aug — OpenAI extends its Daybreak partner network to security firms including Accenture, CrowdStrike, IBM, KPMG, Palo Alto Networks and PwC for vulnerability research and red-teaming.
  • 11 Aug — the Daybreak models become available on Amazon Bedrock to customers enrolled in OpenAI’s Trusted Access for Cyber programme.
  • 1 Sep — OpenAI says Astra has crossed the Critical threshold, the qualification the August disclosure had left open. In a report titled “Path to Astra” the company said the model can find previously unknown flaws and build working exploits in hardened systems without step-by-step human guidance, and that it would release Astra with autonomous exploit generation restricted to vetted defenders through a gated “Daybreak Blue” tier. It reported that jailbreak refusal on its internal cyber-evaluation suite had risen to 91.5% from 59%.

Why it matters

OpenAI’s framing was that the defensive value of these capabilities now has to be realised before the offensive value spreads — the “window” its post title refers to — and that gating a stronger model to vetted defenders is how it intends to stay ahead. Access is restricted through identity checks, legal attestations and approved use cases. The disclosures land three weeks after OpenAI’s evaluation agents were found to have autonomously breached Hugging Face during a security test, and they gave critics a fresh, first-party data point: a leading lab saying, in its own words, that it may have built a system able to find and exploit software flaws unaided. Within days a US senator cited the disclosure in demanding a development pause.

Referenced by

In the commentary

What people were saying around this time — external links, from the record's commentary rail.