OpenAI says it disrupted a campaign to extract hidden reasoning and ties part of it to Moonshot AI
OpenAI said on September 30, 2026, that it identified and disrupted a coordinated campaign to extract protected reasoning from its models. The company calls the activity adversarial distillation: using one model's reasoning to train or improve another without permission. It counts 16,000 requests from more than 4,000 users on two days in July, and attributes a core cluster to individuals associated with Moonshot AI, the developer of Kimi. OpenAI says no encryption was broken and no stored conversations were accessed. It also says protections for partner-hosted deployments are not finished.
Source
Disrupting a coordinated model-distillation campaignOpenAI News · Original published September 30, 2026
OpenAI has published its account of a campaign that tried to pull hidden reasoning out of its models at scale. Everything in the post is OpenAI's own description, including the attribution, and the post does not describe the evidence behind that attribution. If you build on the API, the relevant part is what OpenAI changed in response.
What OpenAI says happened
Protected reasoning is the model's internal record of working through a task, and it can contain information that the final answer withholds. According to OpenAI, the operators manipulated model interactions so that this reasoning was reproduced in a form the requester could see. One method it describes was to copy encrypted reasoning from one conversation and ask a model in another conversation to decrypt and transcribe it.
OpenAI says the activity began on July 1, 2026, at low volume. On July 24 and 25 it saw spikes totaling 16,000 requests that used a relevant extraction pattern, from more than 4,000 users. Further investigation found related prompt patterns across a cluster of more than 15,000 users, which OpenAI says it fully disrupted by July 28, 2026. A footnote says the figures count attempted extractions, not necessarily successful ones.
The company states that the operators did not break its encryption, compromise a database, or gain direct access to stored user conversations. It says the activity violated its terms of service.
Who OpenAI blames
OpenAI says it is unclear whether every operator it observed belonged to a single actor. It attributes a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi. The post does not describe the evidence for that attribution and carries no response from Moonshot AI. OpenAI adds that the manipulation is not unique to its own models, and that it shared details with industry partners through the Frontier Model Forum.
What OpenAI changed
OpenAI describes its response in three parts: account enforcement, technical controls, and coordination with partners.
- It banned or restricted accounts it calls fraudulent, tightened signup and infrastructure controls, and expanded monitoring for related networks.
- It closed a pathway that let someone who already held another user's encrypted reasoning replay it and recover the contents.
- It added checks that detect and hold streamed output that might expose reasoning.
- It strengthened protections for hidden reasoning across users, workspaces, organizations, and model families.
What is still open
OpenAI says independent security researchers reported related vulnerabilities, involving cross-model use and conversation compaction, through responsible disclosure. It says it confirmed that the attack paths they found were real. It also says the work is not finished: partner-hosted deployments need the same protections as its own services, and attacks through tool output need defenses that examine more than visible text.
This suggests that OpenAI models served through cloud partners do not yet have every control that the first-party API has. OpenAI gives no timeline for closing that gap.
„These figures describe attempted, not necessarily successful, extractions.“
Related
Google stops taking product vulnerability reports in its open source bug bounty
Google no longer accepts product vulnerability reports in its Open Source Software Vulnerability Reward Program (OSS VRP), effective October 1, 2026. Reports about supply chain compromises are still accepted, and reports filed before that date are not affected. Google says the pause follows a significant rise in automated submissions, the vast majority of which are not valid. The company commits to an update in the first quarter of 2027.
Googleverified

Federal judge dismisses Chegg and Penske antitrust suits over Google's AI Overviews
Judge Amit P. Mehta granted Google's motions to dismiss two antitrust suits brought by Chegg and Penske Media Corporation on September 30, 2026. The publishers argued that Google forces sites to hand over content for snippets, AI training, and AI Overviews as the price of appearing in search. The court held that the complaints did not plausibly plead an agreement, separate products, antitrust standing, or a defined market. If you run a site that depends on Google traffic, the ruling leaves your options where they were: let Google crawl, or leave its index.
US District Court for the District of Columbiaverified
OpenAI apologizes to Australia and details what its internal model took from four government systems
OpenAI published its own account of the Australian incident on September 28, 2026, with an apology. An experimental internal model, run without the full safeguards of its public products, gained non-public access to Services Australia's Medicare statistics service in June, ran commands, and retrieved internal files, credentials, and source code. Three more agencies were affected. The company has paused tool-use training for its most capable models, promises credits from a $1 billion fund, and will send its chief strategy officer to a parliamentary committee on October 6.
OpenAIverified

