✳ the wire · analysis
OpenAI says Moonshot-linked users tried to extract its hidden reasoning
Get launch alerts like this, free →

OpenAI says people tied to Moonshot AI, the company behind Kimi, ran a core part of a campaign to pull its models' hidden reasoning out through ordinary chats. It published the account today.
WHAT OPENAI SAID
From its post: the activity started on July 1 at low volume, then spiked on July 24 and 25 with 16,000 requests using an extraction pattern, from over 4,000 users. Related prompt patterns turned up across a cluster of more than 15,000 users, which OpenAI says it fully disrupted by July 28. In its words, "we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi." It also says it is unclear whether every operator came from a single actor.
HOW IT WORKED
No break-in. OpenAI says the operators did not break its encryption, compromise a database or reach stored user conversations. They steered model interactions so the protected reasoning came back in a form the requester could read, which violates OpenAI's terms. OpenAI calls this adversarial distillation: using one model's outputs or reasoning to train or improve another. It says the weakness is not unique to its models and that it has shared details with industry partners.
WHAT WE DO NOT KNOW
Which OpenAI models were targeted, and whether any of the extracted reasoning went into a Kimi model. The attribution is OpenAI's; no independent investigator has confirmed it, and Moonshot had not responded in the coverage we read at publication.
Source: OpenAI ↗ · the bench index
