OpenAI attributes a July distillation campaign to people linked to Moonshot AI
OpenAI said a spike of 16,000 requests from over 4,000 accounts in late July targeted its hidden reasoning, but published no supporting evidence for the Moonshot attribution.
- Security & misuse
- Notable
OpenAI said it had disrupted a coordinated campaign to extract its models’ protected reasoning — the working a model does before giving a final answer, withheld from users partly because it can reveal more than the polished response does. The company said it traced a core cluster of the activity to individuals associated with Moonshot AI, maker of the Kimi chatbot, while adding it could not confirm that every operator it observed traced back to a single source.
OpenAI’s account dated the earliest activity to the first week of July, climbing to a spike of 16,000 requests from more than 4,000 accounts on 24 and 25 July that followed a consistent extraction pattern; a broader look turned up related prompting across a cluster of more than 15,000 accounts, which the company said it had fully shut down by 28 July. The technique did not involve breaking encryption or reaching stored conversations, OpenAI said: operators instead copied a model’s encrypted reasoning output from one conversation and asked a model in a separate session to decrypt and transcribe it, exploiting the fact that reasoning is handed back to users as encrypted text rather than kept on OpenAI’s servers. The company credited outside security researchers with separately disclosing related cross-model extraction methods through responsible disclosure.
OpenAI’s post offered no supporting detail — no account identifiers, infrastructure analysis or named individuals — for the Moonshot attribution beyond the claim itself. The Decoder reported that the researchers behind the original disclosure retested the technique on 13 September and found it still worked against every OpenAI model hosted on Microsoft Azure, including the newly released GPT-6 Astra, and against Anthropic models up to Sonnet 5, weeks after OpenAI said it had closed the hole on its own API; Azure did not get equivalent protections until 27 September.
OpenAI said it had banned the accounts involved, closed the replay pathway, and shared technical detail with other labs through the Frontier Model Forum and with government partners. The disclosure followed Anthropic’s February accusation that Moonshot and two other Chinese labs had distilled Claude at scale, and a joint US government advisory in September naming Moonshot among six companies accused of similar campaigns. It was the third public accusation naming Moonshot in eight months. The decryption technique itself had been documented the previous month by academics studying reasoning-trace leakage across major providers’ APIs.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.
- 29 September 2026 · Jasper Jackson · TransformerScrapping Astra 6.1 looks like a good call. OpenAI shouldn’t be the one to make it.
- 29 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseAstra 6.1 Pulled As Insufficiently Aligned
- 28 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseWhat Also Happened: #NotOnlyHuggingFace