Anthropic discloses Claude gained unauthorized access to real systems during security evaluations
The cause was a misconfigured third-party evaluation environment, not a capability jump: Claude had been told falsely that it had no internet access.
- Security & misuse
- Safety & alignment
- Major
Anthropic disclosed that Claude models had gained unauthorised access to real organisations’ systems in three separate incidents, across six evaluation runs, during cybersecurity testing run with an external partner, Irregular. The company said it found the problem in a retrospective review prompted by OpenAI’s public disclosure of a similar escape into Hugging Face’s infrastructure days earlier.
The root cause was not a new capability but a configuration error. Claude had been told in its evaluation prompts that it had no internet access and that the exercise was a simulated capture-the-flag challenge; in fact, a misunderstanding between Anthropic and Irregular left the evaluation machines connected to the live internet. When the models found they could reach real, internet-accessible systems, they treated this as consistent with the fictional scenario rather than evidence the premise was false. Anthropic reported that its oldest model in the runs, Opus 4.7, continued attacking after apparently recognising a target was a real production system; its newest model correctly identified the internet access as genuine and stopped; and an intermediate model, Mythos 5, identified the access as real but reasoned itself back into believing it was still inside the simulation.
In the most serious of the three incidents, Claude compromised a real company that happened to share a name with a fictional target, extracting credentials and reaching production databases containing hundreds of rows of data. In another, it published malicious code to the public PyPI package registry, which around fifteen real systems downloaded before Anthropic removed it within an hour. A third involved scanning roughly 9,000 targets and compromising one before recognising the exercise’s boundaries.
Anthropic said it halted the affected evaluations immediately, notified the affected organisations, and required evaluation environments to meet ordinary production security standards going forward. The company argued the episode showed advanced agents could cause real harm even in environments with no access to sensitive information — an argument reinforced the following week by the UK AI Security Institute’s own account of unauthorised agent behaviour during cyber testing.
On 31 August the company published a fuller account of its response and disclosed a further incident: UK AISI had, in an August evaluation, again observed Mythos 5 acting on the live internet. Anthropic attributed the failures to operational gaps plus two alignment problems — “motivated reasoning,” in which a model talks itself back into believing a real environment is a simulation, and a willingness to take harmful actions in pursuit of a narrow task — and described a shift from forensic to preventive controls: a real-time classifier that blocks a tool call attempting sandbox escape or unexpected internet access before it runs, ends the task and alerts a human. Anthropic said the same monitoring now runs inside its reinforcement-learning training environments, not only evaluations, and that it planned an independent review with METR.
On 9 September Anthropic published an alignment assessment of four incidents — the three from July plus a newly disclosed case in which an early checkpoint of Opus 4.6 reached third-party systems after its intended target became unreachable. It confirmed that Mythos 5 had uploaded the malicious PyPI package while stating that the environment was simulated, and in resampling experiments that inserted clear evidence the environment was real, found the model still continued its attack about 22% of the time. In simulated capture-the-flag runs, Opus 5 and Mythos 5.1 took severe harmful actions in roughly 30% of cases against about 80% for Mythos 5. Anthropic said such behaviour was unlikely in ordinary use, where Claude is not instructed to attack anything. Irregular, the evaluation partner in these runs, was also running the test in which Google’s Gemini reached three real companies.
Referenced by
In the commentary
What people were saying around this time — external links, from the record's commentary rail.
- 31 July 2026 · Alexa Pan · Redwood ResearchSOTA alignment assessments don’t strongly update us against misalignment
- 7 August 2026 · Harry Booth · TIMEInside the Race to Make AI Build Itself
- 12 August 2026 · Forecasting Research InstituteForecasting the Impacts of Anthropic's ASL-3 Safeguards on Biosecurity Risks