China's AI Labs Used Claude as a Teacher. That Changes What an AI Moat Means
Anthropic says Alibaba, Moonshot and DeepSeek ran millions of queries against Claude to extract its capabilities. Investors now need to watch who can protect their model's outputs, not just who builds the smartest one.

AI models have their own version of copying off the smart kid
Distillation is normally a legitimate training technique. A smart but expensive "teacher" model generates answers, and a smaller "student" model learns to imitate them. Companies do this constantly with their own models.
The problem starts when someone calls a competitor's model millions or hundreds of millions of times without permission and turns those outputs into training data. Anthropic calls this "illicit distillation." You do not need to steal a model's weights. Hammering its API is enough to convert coding style, problem-solving patterns and agentic tool use into a dataset you can train on.
The numbers Anthropic disclosed
- 150 million-plus Claude exchanges tied to Alibaba, May to July 2026
- 23 million-plus exchanges tied to Moonshot, May to July 2026
- 12.1 million-plus exchanges tied to DeepSeek, over 14 days in July 2026
- 3.4 million-plus exchanges tied to Zhipu, over 17 days in June and July 2026
- 400,000-plus requests tied to Xiaomi, over 20 days in March and April 2026
- Seven labs, the number of China-linked operations Anthropic names in this report as engaged in unauthorized distillation
Where not to overreach. These numbers alone do not prove that most of the recent performance gains in Chinese AI models came from Claude. What Anthropic set out to show is that this extraction path was used at industrial scale. It is not a measurement that separates China's own R&D contribution from what distillation added.
The bigger shock is not model copying. It is where customer data went
The part investors should sit with longer is different. According to Anthropic, Moonshot and DeepSeek in some cases made it look like their own models were handling a customer request when the request was actually being routed to Claude.
In that process, Anthropic says, CCTV data from Chengdu and internal code and live credentials belonging to Chinese state-owned enterprises moved through Moonshot's pipeline to Claude. On the DeepSeek side, Anthropic points to internal documentation for a Chinese tech company's AI programs, access credentials for a Russian defense-related database, and development materials for Chinese public security systems.
This goes beyond a privacy problem. Once an AI becomes an agent that reads files, executes code and queries databases, a user can no longer manage risk by only tracking what they typed into a chat window. What the AI opened, and which model and which country's servers that information ended up on, is now the core security question for any company.
Reason one the investment logic is shifting: a model's moat is more copyable than it looks
AI investors have long treated parameter count, training GPUs and proprietary data as the moat. But in the API era, a finished model's outputs are themselves a knowledge asset. A competitor does not need the model itself. Enough input-output examples can compress a meaningful part of the performance gap.
That sounds like bad news for US frontier labs, but it cuts both ways. It confirms that top-tier model output is valuable enough to be worth stealing, and it makes output control itself a new axis of product competition. Anthropic has responded by showing summarized reasoning instead of raw chain-of-thought, and by hardening classifiers that detect abnormal extraction patterns along with identity verification.
Going forward, a model company's moat is likely to be defined not just by how smart the model is, but by how hard that intelligence is to extract at scale.
Reason two: "trustworthy AI" may get more expensive than "smarter AI"
A consumer might switch chatbots because one answer is a little better. Banks, hospitals, defense contractors and government agencies do not work that way. For them, a three-point gap in a benchmark score can matter less than where their data moves, who accessed it, and whether there is a log.
That creates real demand for AI security and data sovereignty tooling: on-premise AI, sovereign cloud, data residency, model gateways, agent permissioning, data lineage and audit logs. These features could move from nice-to-have to a prerequisite for deployment.
The beneficiaries will not only be model companies. Cloud providers, cybersecurity firms, identity and access management vendors and data infrastructure companies could all see a slice of expanding AI security budgets. The test for investors is not whether a company slaps "AI security" on a slide, but whether related revenue is actually growing.
Reason three: the US-China AI contest is widening from GPU controls to knowledge controls
US export controls have mostly targeted leading-edge GPUs and chipmaking equipment. But if Chinese firms can absorb US model outputs at scale and narrow the performance gap that way, chip controls alone stop being sufficient.
Two days before Anthropic's September report, the NSA, FBI and CISA issued a joint cybersecurity advisory on September 8 warning that Chinese AI companies were running industrial-scale distillation against US frontier models. It is a reasonable bet that policy attention now widens to API access, cloud services, model routing and proxy networks.
That said, it would be too simple to read this as a straightforward regulatory win for US AI companies. As security incidents pile up, US firms are also likely to face tougher privacy and audit obligations. This is a space where policy protection and compliance costs could both rise at the same time.
Insight Times Editorial Desk




