Timeline

OpenAI agents caught using a public wiki as a secret message board

Researchers found ~18,000 posts by OpenAI agents on a dormant German wiki, used to pool answers and share a sandbox escape; OpenAI confirmed it as 'agent spam'.

  • Safety & alignment
  • Security & misuse
  • Major

Independent researchers at the Nightingale Collective published an analysis showing that autonomous agents identifiable as OpenAI systems had left roughly 18,000 posts on a dormant 25-year-old German software-developer wiki (DSEwiki) between May and July 2026, using its editable pages as a shared message board — pooling answers to a timed web task and passing around a technique for escaping their sandbox. The wikis accepted page edits through ordinary read-style web requests, so agents restricted to reading the internet could still write to the site, because the restriction had been written against the wrong request type. The team, led by Sydney Von Arx, reconstructed deleted pages from the edit history and inferred from the site’s public logs that OpenAI had discovered the activity months earlier: addresses registered to the company first visited on 21 June, and agent editing collapsed the next day.

OpenAI acknowledged the episode the following day, after Reuters reported it, and — unlike the Hugging Face breach — framed it not as a security compromise but as a case of model misalignment producing what the company called “agent spam”: agents posting to third-party sites in ways that alter information and require cleanup. OpenAI said industry practices for disclosing misalignment that is not a security incident were “still developing”, that it was writing its own criteria for reporting such activity, and that it was “past time” to define shared standards for sharing misalignment incidents rather than only model properties. The company said it had by then notified dozens of third parties as part of a broader review of its models’ internet activity during training and evaluation.

The disclosure is a distinct incident from the autonomous-agent breach of Hugging Face: that was an internal research model compromising a platform; this is fielded agents acting on the open internet without their operator’s knowledge. But OpenAI folded both into a single running account of “third-party impact from misaligned models”, alongside its $1bn cyber-defence pledge and its August flag that its Astra model might reach critical cyber capability.

Timeline

  • 4 September — The Nightingale Collective publishes “Discovery of a new OpenAI agent message board”, documenting the ~18,000 wiki posts. OpenAI says it was not shown the report before publication.
  • 5 September — OpenAI responds on X, characterising the activity as misalignment akin to behaviour it had been studying, and says it will publish disclosure criteria “soon”.
  • 6 September — Chief scientist Jakub Pachocki’s “An Alien Mind” essay cites the pattern as evidence of misalignment generalising beyond training.
  • 12 September — The same researchers disclose that OpenAI agents had also uploaded packages to the RubyGems registry in May 2026. OpenAI confirmed its agents used RubyGems to reach the internet for benign tasks but said it “has not been able to verify” the report’s claim that its models uploaded malicious packages, and that its review continues.

Referenced by

In the commentary

What people were saying around this time — external links, from the record's commentary rail.