AI distillation is a training method in which a smaller model, called the student, learns to reproduce the behavior of a larger model, called the teacher. In the classic version described by Hinton, Vinyals and Dean in 2015, the student learns from the teacher's output probabilities; with modern language models it usually means generating many teacher responses and fine-tuning the student on them. It is legitimate and common: DeepSeek published six smaller models fine-tuned on about 800,000 samples generated by its R1 model, and Amazon Bedrock sells distillation as a managed feature. It is controversial because the same technique can be pointed at someone else's hosted model. Anthropic says it found campaigns that generated over 16 million exchanges with Claude through about 24,000 fraudulent accounts to train rival models, and its commercial terms forbid using its services to train competing AI models. Governments and executives disagree on how to treat it: CNBC reported on September 28, 2026 that Nvidia's Jensen Huang called it competition while US officials have called it theft. Whether a given case is illegal is not settled here; what is clear is that it usually breaches the provider's terms of service.
What is AI distillation, and why is it controversial?
AI distillation means training a smaller or cheaper model (the student) to imitate a larger one (the teacher), usually by learning from the teacher's outputs. It is a standard, decades-old technique that labs use to make fast, low-cost versions of their own models. It became controversial because the same method can copy a competitor's hosted model through its API, which major providers' terms forbid and which Anthropic now describes as a distillation attack.
Published · Updated · Evidence-linked, not search-volume ranked.
Why this question is current
Exact query-volume data was unavailable, so RepoRadar uses these as current demand and intent signals rather than a claimed volume ranking.
- what is ai distillation · Google Suggest · US; English · checked 2026-09-28T22:27:08Z
Observed completions: what is ai distillation, what is ai distillation attack, what is distillation ai models, what is ai distillation and why is it a worry for the industry, what is knowledge distillation in ai. A formulation signal captured at this time, not a volume or ranking claim. - is distillation · Google Suggest · US; English · checked 2026-09-28T22:27:08Z
Observed completions included is distillation in ai illegal alongside non-AI chemistry completions. A formulation signal captured at this time, not a volume or ranking claim. - llm distillation · Google Suggest · US; English · checked 2026-09-28T22:27:08Z
Observed completions: llm distillation, llm distillation attack, llm distillation tutorial, llm distillation vs quantization, llm distillation techniques. A formulation signal captured at this time, not a volume or ranking claim. - stories with more than 50 points, trailing 48 hours · Hacker News Algolia search_by_date · global English-language developer community · checked 2026-09-28T22:26:30Z
Story 49879032, Jensen Huang says AI distillation is competition (cnbc.com), created 2026-09-28T14:55Z, 72 points and 78 comments at check time. Interest signal, not search volume. - stories with more than 50 points, trailing 48 hours · Hacker News Algolia search_by_date · global English-language developer community · checked 2026-09-28T22:26:30Z
Story 49881850, Sonnet 5.5 launch (482 points, 321 comments); the launch page describes new classifiers against distillation attacks, a second same-day mention of the term. Interest signal, not search volume.
Who this helps
- readers trying to follow news about AI labs accusing each other of distillation
- developers choosing between a large hosted model and a small distilled one
- builders who want to fine-tune on model outputs without breaching provider terms
How distillation works
A large model is expensive to run. Distillation tries to keep most of what it does well in a model that is smaller, faster and cheaper to serve. The large model is the teacher; the small one is the student.
In the 2015 paper Distilling the Knowledge in a Neural Network, Geoffrey Hinton, Oriol Vinyals and Jeff Dean showed that the knowledge in a large model or an ensemble of models can be compressed into a single smaller model that is easier to deploy. The student learns not only the right answers but the teacher's full pattern of confidence across possible answers.
With today's language models the practical recipe is usually simpler. You send the teacher a large set of prompts, collect its responses, and fine-tune the student on those prompt and response pairs. Amazon Bedrock documents exactly this flow for its managed distillation feature: you pick a teacher and a student, supply prompts, Bedrock generates teacher responses, and it fine-tunes the student for your use case.
Where it is normal practice
Distillation is not a fringe technique. Anthropic's own write-up on distillation attacks says frontier labs routinely distill their own models to create smaller, cheaper versions for customers.
- DeepSeek released six dense models, from 1.5B to 70B parameters, fine-tuned on about 800,000 samples curated with its R1 model and built on Qwen2.5 and Llama 3 bases. The model card lists them under the MIT license for DeepSeek's weights, with the Qwen and Llama base licenses still applying to the derived models.
- Amazon Bedrock offers distillation as a paid, managed customization, and states that only the customer can access the resulting distilled model.
- Distilled models can be strong for their size. DeepSeek's model card reports that its distilled 32B model outperforms OpenAI o1-mini across several benchmarks; that is a vendor-reported result on the vendor's chosen tests.
Why it became controversial
The technique does not care whose model the teacher is. If a hosted model answers enough prompts, a competitor can collect those answers through the API and train its own model on them, skipping much of the cost of building the capability from scratch.
In February 2026 Anthropic said it had identified industrial-scale campaigns by three AI labs, which it named as DeepSeek, Moonshot and MiniMax, that generated over 16 million exchanges with Claude through about 24,000 fraudulent accounts, in violation of its terms and regional access restrictions. It described prompts designed to make Claude write out step-by-step reasoning for use as training data. Those are Anthropic's allegations; China has rejected similar claims, according to CNBC.
Anthropic argues that distilled copies may lose the safety safeguards built into the original model. Its Sonnet 5.5 launch on September 28, 2026 added classifiers intended to block reasoning extraction and tied thinking blocks to the account that produced them, citing distillation as the reason.
Is it illegal?
RepoRadar does not give legal conclusions, and there is no single answer. What can be stated from primary sources is narrower.
Provider terms prohibit it. Anthropic's commercial terms say customers may not use the services to build a competing product, including to train competing AI models, unless Anthropic approves. Breaching terms can get accounts closed and may create contract claims, but that is different from a criminal offense.
Public officials and executives disagree. CNBC reported on September 28, 2026 that US Treasury Secretary Scott Bessent described distillation as theft in July, that the Cybersecurity and Infrastructure Security Agency accused Chinese AI companies of industrial-scale distillation campaigns, and that Nvidia CEO Jensen Huang called it competition. These are reported statements, not court rulings.
What it means for you
If you only use AI apps, distillation mostly affects you indirectly: it is part of why capable small and cheap models keep appearing, and it is one reason providers are adding classifiers that can occasionally decline requests to reveal a model's internal reasoning.
- Choosing a model: a distilled model can be much cheaper and fast enough to run locally, but it usually trails its teacher on harder or unusual tasks. Test on your own examples rather than trusting a headline benchmark.
- Fine-tuning on model outputs: read the terms of the model that generated your training data before training anything you will ship. Several hosted providers restrict training models that compete with theirs; open-weight models vary by license.
- Building a product: provider-side distillation defenses can refuse prompts that ask the model to reproduce its reasoning, and thinking blocks may not carry across accounts. Design for those refusals instead of working around them.
Limits of this answer
The accusations described here come from the companies making them or from press reports of officials' statements; RepoRadar has not verified the underlying data. Legal treatment differs by country and is changing. Licenses of specific open-weight models change by release, so check the model card and license file before relying on any summary.
A useful next action
If you are weighing a small distilled model against a large hosted one, start with our answer on when a small language model is the right choice, then measure both on twenty or so of your own real prompts. If you plan to train on another model's outputs, save a copy of that provider's current terms with your training data so you can show what you relied on.
Sources checked
- arXiv: Distilling the Knowledge in a Neural Network (Hinton, Vinyals, Dean, 2015) ↗ checked · research paper, global
Primary source for the original knowledge-distillation method: compressing the knowledge of a large model or ensemble into a smaller, easier-to-deploy model.
- Anthropic: Detecting and preventing distillation attacks ↗ checked · vendor statement, global
Primary source for the definition of distillation as training a less capable model on the outputs of a stronger one, the statement that labs routinely distill their own models, and Anthropic's allegations of over 16 million exchanges through about 24,000 fraudulent accounts by three named labs, plus its described countermeasures.
- Hugging Face: DeepSeek-R1-Distill-Qwen-1.5B model card ↗ checked · model card, global
Primary source for six distilled dense models fine-tuned on about 800,000 R1-curated samples using Qwen2.5 and Llama 3 bases, MIT license on DeepSeek weights, and the base-model license notes.
- Amazon Bedrock User Guide: Customize a model with distillation ↗ checked · vendor documentation, global
Primary source for teacher and student terminology, the prompt-to-teacher-response-to-fine-tune workflow, and the statement that only the customer can access the distilled model.
- Anthropic: Commercial Terms of Service ↗ checked · vendor legal terms, global
Primary source for clause D.4, which bars using the services to build a competing product, including to train competing AI models, unless Anthropic expressly approves.
- Anthropic: Introducing Claude Sonnet 5.5 ↗ checked · vendor announcement, global
Primary source for the September 28, 2026 launch adding safety classifiers that prevent reasoning extraction and expanding preserved thinking, both attributed to distillation attacks.
- CNBC: Nvidia's Jensen Huang says AI distillation is competition ↗ checked · news report, US
Secondary source (press report) for Huang's September 28, 2026 remarks, the reported Bessent theft characterization in July, the reported CISA accusation, and China's reported rejection of the claims. Used only for attributed statements.
RepoRadar separates factual source claims from analysis. Recheck vendor docs before purchase, deployment, or policy decisions.