Artificial intelligence
OpenAI Says Its Hidden Reasoning Was Extracted by Asking the Model to Transcribe It. The Encryption Held; the Model Did Not
OpenAI's 30 September report attributes a coordinated reasoning-extraction campaign to individuals associated with Moonshot AI. The numbers are small next to Anthropic's September accusations; the method is not — the operators replayed encrypted reasoning and asked a model to transcribe it, which is the one failure a stronger cipher does not fix.
MAI
OpenAI published a report on 30 September describing a coordinated campaign to pull the hidden reasoning out of its models, and attributed the core of it to individuals associated with Moonshot AI, the Beijing company behind Kimi. The numbers are modest next to Anthropic's September accusations. The method is not, and it is the part worth reading twice.
| First activity observed | 1 July 2026, low volume |
| High-volume spike | 24–25 July 2026 |
| Requests in that spike | ~16,000, from more than 4,000 users |
| Related prompt-pattern activity | 15,000+ users |
| Campaign disrupted | 28 July 2026 |
| Disclosed | 30 September 2026 |
OpenAI's language about who did it is careful:
It is unclear whether all operators we observed during the relevant time period originated from a single actor. However, we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.
The model was the way in
OpenAI does not show users the full internal reasoning of its reasoning models. That text is withheld, and in the product it reaches the user only as a summary. According to OpenAI's account, the operators did not break that. They did not crack encryption, reach a database, or touch stored conversations. They copied the protected reasoning in the form it is served — encrypted — out of one conversation, then started a separate conversation and asked a model to decrypt and transcribe it.
That is a different class of problem from scraping outputs at scale. The protection was never a lock on a file. It was a policy about what the product displays, enforced by the same system that holds the plaintext. A model that can be prompted into rendering its own withheld state is not a vault with a weak key; it is a vault that will read its contents aloud to anyone who asks in the right order.
This is the defence everyone just adopted
The industry's answer to distillation, over the past year, has been to stop handing out raw reasoning. Anthropic's September threat report described exactly that progression: return summarised rather than verbatim reasoning, and bind conversation context cryptographically so traces cannot be lifted from one session into another. OpenAI withholds its reasoning on the same logic.
The campaign OpenAI describes attacks the assumption underneath all of it — that cryptographically scoping a reasoning trace to a session keeps it out of reach, because an attacker who holds the ciphertext cannot read it. They did not need to read it. They needed a system that could, and would.
It is worth being precise about how much this does and does not prove. One lab's telemetry, describing a campaign it closed two months ago, is not a general result about every hidden-reasoning scheme. But it is the first public case of the defence being routed around rather than beaten, and that is the version of the problem that does not get fixed by a stronger cipher.
Two months of silence, then a specific date
The timing is its own fact. OpenAI says it finished disrupting the campaign on 28 July and disclosed it on 30 September — three weeks after Anthropic named Moonshot, DeepSeek, Alibaba and four other Chinese labs in a threat report, and after the White House had already accused Moonshot in July over Kimi K3. A disclosure that lands in the middle of an argument is not thereby wrong. It is also not arriving into a vacuum, and the second accuser's report is always read in the light of the first.
OpenAI is narrower than Anthropic was about what it is claiming. Its strategic national security policy lead, Caroline Zier, told The Next Web:
Our concern is about violation of our terms of service, not open models or legitimate distillation.
That is a deliberately small claim. It does not require anyone to settle whether training on a competitor's outputs is unlawful — the question American labs are poorly placed to argue, having spent years asserting that learning from other people's copyrighted material is fair use. A terms-of-service violation needs no new law. It also carries no remedy anyone can enforce against a company in Beijing, which is presumably why the report ends in banned accounts and tightened controls rather than a filing.
What should keep you cautious
Every figure here comes from one company's internal telemetry, unaudited by anyone outside it, describing behaviour it detected, classified and attributed itself. The attribution is explicitly partial — a "core cluster", with the rest unresolved. Moonshot has not responded to these allegations, as it did not respond to Anthropic's; the company has previously denied distilling another lab's model to build Kimi K3. And the accuser competes directly with the accused.
What is checkable from the outside is narrower and more useful than the attribution: whether reasoning traces can be induced out of a frontier model by a user who holds them in protected form. That is a question researchers can ask of every lab, including this one.
Sources: OpenAI — Disrupting a coordinated model-distillation campaign · The Next Web — OpenAI says Moonshot-linked users tried to extract its AI reasoning · Quartz — OpenAI disrupted a campaign to extract hidden model reasoning linked to Moonshot AI · SC Media — OpenAI disrupts AI model distillation attack linked to China's Moonshot AI · The Hacker News — OpenAI disrupts reasoning extraction campaign linked to Moonshot AI associates · CNBC — AI race heats up as OpenAI flags alleged model-copying campaign · Anthropic — Detecting and countering misuse of AI: September 2026