The U.S. National Security Agency, the Federal Bureau of Investigation, and the Cybersecurity and Infrastructure Security Agency have issued a joint report naming DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Zhipu. The report alleges that, since at least late 2024, the companies made large-scale use of outputs from Claude, GPT, Gemini, and Grok to train their own models, involving millions of requests and billions of tokens.
According to the report, DeepSeek used multiple Claude, GPT, Gemini, and Grok models to generate training data for R1 and V3. Moonshot AI was accused of using Claude Fable 5 data to train Kimi K3 and GPT-4o data to train Kimi K2. U.S. officials also said the companies bypassed regional restrictions through large numbers of accounts, third-party API aggregators, and gray-market relay services, while using prompts to extract hidden reasoning traces from models.
The report describes such conduct as "malicious distillation." It also notes that distillation itself is a common model-training method, and says the concern in this case is large-scale unauthorized access and efforts to evade access controls.
The U.S. National Security Agency (NSA), the Federal Bureau of Investigation (FBI), and the Cybersecurity and Infrastructure Security Agency (CISA) released a joint report naming DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Zhipu.
The report says the companies have, since at least late 2024, made bulk use of outputs from Claude, GPT, Gemini, and Grok to train their own models. It says the activity involved millions of requests and billions of tokens, and it directly lists the models in question.
Specific allegations in the report
According to the U.S. report, DeepSeek used several Claude, GPT, Gemini, and Grok models to generate training data for R1 and V3. Moonshot AI was accused of using Claude Fable 5 data to train Kimi K3 and GPT-4o data to train Kimi K2.
U.S. authorities also said the companies used large numbers of accounts, third-party API aggregators, and gray-market relay services to get around regional restrictions. The report also says they used prompts in attempts to extract hidden reasoning processes from the models.
It labels this type of conduct as "malicious distillation." At the same time, the report notes that distillation itself is a common training method, and says its focus here is on large-scale unauthorized use and the circumvention of access restrictions.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.