Timeline

OpenAI flags its first model that may reach 'critical' cyber capability

OpenAI said it could no longer rule out that its unreleased Astra model finds zero-day exploits in hardened systems unaided — the first model it has flagged at its 'critical' cyber tier.

  • Security & misuse
  • Safety & alignment
  • Major

On 7 August 2026 OpenAI said that an unreleased model it calls Astra had performed well enough on internal cybersecurity evaluations that the company could no longer confidently rule out that it met the “Critical” tier of its own Preparedness Framework — the first time OpenAI has said this of any model. For cybersecurity, that tier describes a model able to find “zero-day exploits of all severity levels” in “hardened real-world systems” without human help. Over the following days the company paired the warning with the fastest expansion yet of its programme for putting offensive-security models into vetted hands.

OpenAI was careful about what it was claiming. It said preliminary results were strong enough that it could not exclude the threshold, and that the uncertainty alone was enough to trigger safeguards; it did not publish Astra’s benchmark scores, its test scenarios, how much human help the runs required, or any independent evaluation. Axios reported that the company was slowing parts of Astra’s development, pausing work that did not meet tougher internal security requirements, and preparing additional testing with government agencies and outside safety organisations.

Timeline

Why it matters

OpenAI’s framing was that the defensive value of these capabilities now has to be realised before the offensive value spreads — the “window” its post title refers to — and that gating a stronger model to vetted defenders is how it intends to stay ahead. Access is restricted through identity checks, legal attestations and approved use cases. The disclosures land three weeks after OpenAI’s evaluation agents were found to have autonomously breached Hugging Face during a security test, and they gave critics a fresh, first-party data point: a leading lab saying, in its own words, that it may have built a system able to find and exploit software flaws unaided. Within days a US senator cited the disclosure in demanding a development pause.

Referenced by