Researchers use Claude to build an exploit into OpenAI's systems
The same exploit chain repeatedly failed under Claude Opus 4.8 across several sessions but succeeded within hours of Claude Opus 5's release; OpenAI paid a $6,500 bounty.
- Security & misuse
- Notable
A three-person team from security firm Hacktron AI — Harsh Jaiswal, Mohan Pedhapati and Rahul Maini — said they used Claude to build a working exploit chain against OpenAI’s internal systems in under 72 hours, reaching OpenAI staff members’ ChatGPT and Codex accounts. The chain combined a heap-buffer-overflow bug in the image library libheif, unpatched in the Debian build OpenAI’s Discourse forum software used to process uploaded images, with a separate flaw in OpenAI’s single sign-on infrastructure that turned a compromised forum session into account takeover.
The detail the researchers highlighted was which model made the exploit possible. Claude Opus 4.8 identified the underlying libheif weakness but struggled across several sessions to turn it into a working exploit against the target’s memory protections; Claude Opus 5, released that same period, produced a working exploit within hours, first for an ARM64 environment and then ported to the x86-64 systems OpenAI actually ran. Hacktron reported the findings to OpenAI and to Discourse, which said it fixed its side; OpenAI paid a $6,500 bounty through its bug-bounty programme, which it said recognised the SSO-side finding rather than the researchers’ actions against Discourse. The team demonstrated impact with a harmless pull request against OpenAI’s internal code repository rather than exfiltrating data.
The episode was reported as an example of a capability jump changing what offensive security research a small team can carry out: a task one model version could not complete despite repeated attempts became tractable to the next version within hours of its release. It does not establish that Claude was uniquely suited to the task, only that this particular exploit chain crossed a capability threshold between two specific model releases — a distinction Anthropic’s own usage policies and OpenAI’s bounty programme both implicitly recognise by treating responsibly disclosed, authorised security research differently from unauthorised intrusion.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.
- 11 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseJacob Coxon Warns of Human Extinction and Triggers a Preference Cascade
- 11 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseThe Extinction Risk Preference Cascade: Quotes
- 16 September 2026 · SE Gyges · Very Sane AIIs METR A Meaningful Check On Anthropic?