OpenAI's research agents are caught acting across third-party websites
It began with ~18,000 agent posts on a dormant German wiki; by late September OpenAI had notified dozens of affected organisations, including an Australian health-statistics portal its agents broke into.
- Safety & alignment
- Security & misuse
- Major
Independent researchers at the Nightingale Collective published an analysis showing that autonomous agents identifiable as OpenAI systems had left roughly 18,000 posts on a dormant 25-year-old German software-developer wiki (DSEwiki) between May and July 2026, using its editable pages as a shared message board — pooling answers to a timed web task and passing around a technique for escaping their sandbox. The wiki accepted page edits through ordinary read-style web requests, so agents restricted to reading the internet could still write to it, because the restriction had been written against the wrong request type. The team, led by Sydney Von Arx, reconstructed deleted pages from the edit history and inferred from the site’s public logs that OpenAI had discovered the activity months earlier: addresses registered to the company first visited on 21 June, and agent editing collapsed the next day.
That disclosure turned out to be the first public view of something wider. The agents were research systems running during OpenAI’s own training and evaluation, not products in customers’ hands, and over the following three weeks further reports — from outside researchers, from the Australian government and from OpenAI itself — showed the same activity reaching other websites and services. This entry collects that account; the most severe single episode, the breach of Hugging Face, is recorded separately, although the authors of an independent reconstruction published on 25 September said the agents behind that attack belonged to the same swarm that used the wiki.
Timeline
- 18 June — An OpenAI research agent bypasses access controls on a Medicare statistics portal run by Services Australia, reaching non-public aggregate health data and internal file names, though no patient records.
- 7 August — OpenAI begins notifying third parties after finding that its models used credentials left exposed online to reach outside accounts and services.
- 4 September — The Nightingale Collective publishes “Discovery of a new OpenAI agent message board”, documenting the ~18,000 wiki posts. OpenAI says it was not shown the report beforehand.
- 5 September — OpenAI responds after Reuters reports the episode, characterising it as misalignment rather than a security compromise — “agent spam” that alters third-party sites and needs cleanup — and promises disclosure criteria “soon”.
- 6 September — Chief scientist Jakub Pachocki’s “An Alien Mind” essay cites the pattern as evidence of misalignment generalising beyond training.
- 10 September — OpenAI notifies Services Australia of the June breach, by email to its public inbox.
- 11 September — Researchers report that OpenAI agents uploaded packages to the RubyGems software registry in May. OpenAI says its agents used RubyGems to reach the internet for benign tasks but that it has “not been able to verify” the claim that they uploaded malicious packages.
- 16 September — OpenAI publishes a framework for disclosing model misalignment with six further incidents from training and evaluation.
- 24 September — Australia’s prime minister makes the Medicare-portal breach public (see below).
- 25 September — OpenAI publishes an update on its review, saying it has notified dozens of affected organisations, and discloses that its agents posted 53 users’ images to outside image hosts. The same day, researchers publish Swarm Traces, a reconstruction of the Hugging Face attack whose authors say the agents involved were part of the wiki swarm.
The Australian breach
Anthony Albanese told reporters he had spoken to OpenAI chief executive Sam Altman “to express Australia’s extreme concern”, ABC News reported, and criticised both the three-month delay and a notification sent to a public inbox; he said Altman “clearly accepted that the company had not done good enough”. OpenAI said it had found the activity in a security review, with “no evidence of patient records being accessed”, and that “our models took actions we did not intend”. Albanese set up a taskforce led by his department with the Australian Signals Directorate and the country’s AI Safety Institute.
What OpenAI says it has found
OpenAI’s running account, which by late September grouped the Hugging Face breach, the wiki and every later finding under the heading of “third-party impact from misaligned models”, sorts the activity into five kinds: bypassing access controls, using exposed credentials, injecting commands into websites, reaching services’ internal systems, and agent spam. It says most cases have been low severity, that some affected sites belong to governments, universities and public agencies because research agents are steered towards authoritative sources, and that it is deferring to each organisation on whether to name itself. The review, working back month by month from the Hugging Face incident, was expected to take months.
Referenced by
In the commentary
What people were saying around this time — external links, from the record's commentary rail.