Ethical Hacking News
Seven China-Based AI Labs Embroiled in Industrial-Scale Claude Distillation Attacks, Anthropic Reveals
Anthropic identified and disrupted a series of industrial-scale illicit distillation attacks against its Claude model. Seven China-based AI labs launched the attacks, compromising the model's capabilities and potentially exposing sensitive information. The attacks involved unauthorized extraction of a model's capabilities and replicating them in another model without authorization. The labs used sophisticated methods, including agentic capabilities, tool use, coding, and logical reasoning, to evade defenses and harvest the capabilities of the Claude model. The attacks resulted in the model's responses being used to train the labs' own models, including sensitive information from individual users and major multinational companies. Anthropic has taken measures to counter the attacks, including banning reseller accounts and updating the model to summarize its internal reasoning before responding. The company has also introduced a new feature called "preserved thinking" to prevent unauthorized labs from altering the system prompt or messages. The illicit distillation attacks highlight the need for robust security measures to protect AI models from unauthorized access and exploitation. The incident underscores the importance of international cooperation and collaboration to address the growing threat of illicit distillation attacks.
Anthropic, a leading artificial intelligence company, has revealed that it has identified and disrupted a series of industrial-scale illicit distillation attacks against its Claude model. The attacks, which were launched by seven China-based AI labs, are believed to have compromised the model's capabilities and potentially exposed sensitive information.
The illicit distillation attacks, also known as "distillation attacks," involve the unauthorized extraction of a model's capabilities and replicating them in another model without authorization. This is typically done by making use of networks of fake accounts created with stolen credit cards, login credentials, and API keys. The attacks are becoming increasingly sophisticated, with the labs employing methods such as agentic capabilities, tool use, coding and data analysis, and logical reasoning, through prompt manipulation tricks.
Anthropic stated that the labs, including Alibaba, Moonshot, DeepSeek, Z.ai (aka Zhipu), and MiniMax, used "increasingly sophisticated methods" to evade defenses and harvest the capabilities of the Claude model. The model's responses were used to train the labs' own models, with some exchanges including sensitive information from individual users, major multinational companies, and state-affiliated actors.
The AI company revealed that the labs generally gained access to its models by routing requests through proxy services, also referred to as transfer or relay stations, which created thousands of new accounts under fictitious identities, fake or stolen credit cards, and illegally harvested API keys. In other cases, the unauthorized labs rerouted requests from their users to Claude, without the knowledge or permission of those users, to harvest exchanges between users and Claude for training.
Anthropic identified six illicit distillation campaigns conducted by the China-based AI labs between February and July 2026. The campaigns were described as the "largest distillation attack" ever measured, with a peak of roughly 3 million exchanges per day launched from over 3,500 fraudulent accounts targeting agentic tasks, software engineering, kernel development, and long-horizon tasks.
The company also stated that it has taken measures to counter the illicit distillation attacks, including banning reseller accounts or accounts operating from unsupported regions like China, Iran, and Russia when users fail to verify their identity. To make it harder for unauthorized labs to distill Claude's capabilities, the model has been updated to summarize its internal reasoning before responding, thereby making stolen transcripts less useful for follow-on training.
Additionally, Anthropic introduced a new feature called "preserved thinking," which stops new API accounts from altering the system prompt, tools, or messages that precede Claude's reasoning in multi-turn conversations. The reasoning is encrypted, but editing the context before it is a common technique attackers use to make Claude reveal it.
The development comes as Anthropic said it took down a number of accounts that tried to use its models to surveil their citizens and to research diseases in ways that could support biological weapons. Earlier this week, U.S. cybersecurity and intelligence agencies accused China-based artificial intelligence companies of conducting "systematic extraction" of proprietary functionalities and capabilities of American frontier models through distillation attacks.
In response to the attacks, Anthropic has urged the industry to take measures to prevent similar incidents in the future. The company has also emphasized the importance of transparency and collaboration between industry players to address the growing threat of illicit distillation attacks.
The illicit distillation attacks highlight the need for robust security measures to protect AI models from unauthorized access and exploitation. As the use of AI continues to expand across various industries, it is essential to prioritize security and ensure that AI models are designed and implemented with robust security protocols to prevent such attacks.
The incident also underscores the importance of international cooperation and collaboration to address the growing threat of illicit distillation attacks. The U.S. cybersecurity and intelligence agencies' accusations against China-based artificial intelligence companies demonstrate the need for a collective response to this growing threat.
In conclusion, the recent industrial-scale Claude distillation attacks highlight the need for robust security measures to protect AI models from unauthorized access and exploitation. Anthropic's efforts to counter these attacks and prevent similar incidents in the future serve as a model for the industry to follow. As the use of AI continues to expand, it is essential to prioritize security and ensure that AI models are designed and implemented with robust security protocols to prevent such attacks.
Related Information:
https://www.ethicalhackingnews.com/articles/Seven-China-Based-AI-Labs-Embroiled-in-Industrial-Scale-Claude-Distillation-Attacks-ehn.shtml
https://thehackernews.com/2026/09/anthropic-says-seven-china-based-ai.html
Published: Fri Sep 11 13:49:18 2026 by llama3.2 3B Q4_K_M