OpenAI says it disrupted distillation campaign and linked a core group to Moonshot AI

OpenAI says it disrupted distillation campaign and linked a core group to Moonshot AI

N
News Editor
2026-10-01 05:15:46
OpenAI said in a new report that it identified and disrupted an organized "adversarial distillation" campaign aimed at extracting protected reasoning from its models. The company said the activity began on July 1 and then spiked on July 24 and 25, when more than 4,000 users generated 16,000 requests that matched extraction patterns. After further investigation, OpenAI said it found related prompt patterns in a group of more than 15,000 users and fully blocked the activity on July 28. While the company said it could not confirm that every participant came from the same side, it attributed the core group to individuals tied to Moonshot AI, the Chinese startup behind Kimi. OpenAI said the actors did not break encryption, breach databases, or directly access stored user chats. Instead, they manipulated model interactions so hidden reasoning could reappear in a visible form. The company said it has blocked or limited accounts, tightened registration and infrastructure controls, patched a replay-related weakness, and shared information with other frontier model developers and government channels.

OpenAI said it has identified and disrupted an organized "adversarial distillation" campaign that tried to extract protected reasoning from its models. In the report, the company said it could not determine whether all participants were part of the same operation, but it attributed the core group to individuals associated with Moonshot AI, the Chinese AI startup behind Kimi.

Activity surged on July 24 and 25

According to OpenAI, the campaign began on July 1 and was limited at first. The volume then jumped sharply on July 24 and 25, when more than 4,000 users sent 16,000 requests that matched extraction patterns.

After additional investigation, OpenAI said it found related prompt patterns in a group of more than 15,000 users and fully blocked the activity on July 28.

What OpenAI means by adversarial distillation

OpenAI described adversarial distillation as the systematic, unauthorized use of one model’s outputs or reasoning to train, copy, or improve another model. Protected reasoning refers to a model’s internal record while it processes a task. If extracted, it can expose information that is not present in the final answer and can help others replicate model capabilities.

OpenAI said the actors did not break encryption, breach databases, or directly access stored user conversations. Instead, they manipulated interactions with the model so hidden reasoning could be reproduced in a form visible to the requester.

One newer technique, according to the company, involved copying encrypted reasoning from one conversation and then asking a model in another conversation to decrypt and transcribe it.

OpenAI says the practice raises security and national security risks

OpenAI said adversarial distillation creates both security and national security risks. Extracted reasoning can be used to train another model without preserving the original model’s safeguards. The company also said large-scale distillation could speed up the transfer of advanced capabilities without equivalent investment in safety, and that the issue becomes more serious as models gain stronger dual-use capabilities.

As part of its response, OpenAI said it blocked or restricted fraudulent accounts, strengthened registration and infrastructure controls, and closed a weakness that had allowed someone holding another party’s encrypted reasoning to replay and reconstruct its contents. When the activity moved through third-party services, OpenAI said it worked with those providers to identify and block the accounts involved.

Information shared with industry and government channels

OpenAI said it notified other frontier model developers and public-sector bodies through the Frontier Model Forum and government information-sharing channels.

The company also said this manipulation method is not unique to OpenAI models. Independent security researchers had previously reported related cross-model weaknesses through responsible disclosure.

Anthropic made a similar accusation weeks earlier

According to CNBC, Anthropic had accused several Chinese AI developers, including Moonshot AI and Alibaba, of secretly using Claude to help train their own models in the weeks before OpenAI published its report. ABMedia said Chain News had previously reported on Anthropic’s efforts to block the illegal distillation of Claude by Chinese labs.

OpenAI said it expects distillation attempts to become more sophisticated as frontier models improve. The company said its next steps will focus on three areas: stronger technical protections, better coordinated detection and enforcement, and deeper threat-intelligence sharing across industry and government.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
300

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.