OpenAI disrupts a campaign to extract its model reasoning, linked to Moonshot AI
OpenAI says a coordinated distillation campaign tried to reproduce its protected reasoning at scale. A core cluster is attributed to individuals associated with Moonshot AI.
OpenAI said on Wednesday, September 30, that it identified and disrupted a coordinated distillation campaign designed to extract protected reasoning from its models. The Hacker News reports that a "core cluster" of the activity, going back to the first week of July, is attributed to individuals associated with Moonshot AI, a Chinese AI company based in Beijing. OpenAI did not publish technical evidence for that attribution.
What happened
According to OpenAI's account as reported by The Hacker News:
- The activity began on July 1 at low volume and spiked on July 24 and 25 to 16,000 attempted requests using an extraction pattern, from more than 4,000 users.
- Related prompt-pattern activity appeared across more than 15,000 users.
- OpenAI fully disrupted the campaign on July 28.
The operators did not break encryption, compromise a database or reach stored conversations. Instead, they manipulated model interactions so that protected reasoning appeared in a form visible to the requester, at scale, which violated the terms of service. OpenAI calls this adversarial distillation: using one model's outputs to train or improve another.
What OpenAI did
It banned the fraudulent accounts, deployed more mitigations, closed a pathway that let someone who already held another user's encrypted reasoning replay it and recover the content, and added checks that hold streamed output that might expose reasoning. OpenAI says extracted reasoning could train another model without the safeguards of the original, and that distillation can speed up the transfer of advanced capabilities without the same safety investment.
Why it matters
Reasoning traces show how a model works through a problem, so they are valuable and hard to protect. The Hacker News adds that this is not the first Moonshot accusation: last month Anthropic accused the company of relaying customer requests to Claude and showing the answers as if from Kimi. An August study found that encrypted reasoning traces from Claude, Gemini and GPT could be swapped across sessions and models within a provider, which can help bypass anti-distillation defenses.
Dany's take
Stolen chain-of-thought is the new IP heist. What stands out is how ordinary it looked: no hack, just clever use of the product at scale. That makes it hard to stop and easy to deny. I would treat the attribution as OpenAI's claim, since no evidence was shown, and wait to hear Moonshot's response.
Source: OpenAI, Disrupting a coordinated model distillation campaign. Details via The Hacker News.
Source: openai.com