Model Distillation Explained: Why AI Labs Guard Against It
Table of Contents

Model Distillation Explained: Why AI Labs Guard Against It

Model distillation explained: Claude Fable 5 routes distillation requests elsewhere. What distillation is, why labs block it, and how to teach it to your kid.

Model distillation is a technique for training a small model to imitate a large one, transferring most of the capability into something much cheaper to run. It’s one of the most useful ideas in machine learning. It’s also, according to Anthropic’s own documentation, one of three request categories serious enough that Claude Fable 5 hands them off to a different model rather than answering.

The other two categories are cybersecurity and biology/chemistry. Distillation sits next to bioweapons risk on that list. Understanding why is a short and genuinely interesting lesson in how the AI industry actually works, and it turns out to be a perfect vehicle for teaching kids the difference between learning from someone and copying them.

Key Takeaways

  • Distillation trains a “student” model on the outputs of a “teacher” model. Hinton, Vinyals, and Dean’s 2015 paper showed you can “compress the knowledge in an ensemble into a single model which is much easier to deploy.”
  • Anthropic’s June 9, 2026 announcement states that when Fable 5’s classifiers detect a request related to cybersecurity, biology and chemistry, or distillation, the response is automatically handled by a more restricted model instead. More than 95% of Fable sessions involve no fallback at all.
  • The reason labs guard against it is commercial: training a frontier model costs enormous sums, and distilling one can reproduce much of its behavior for a fraction of that.
  • Distillation used legitimately is everywhere and makes AI affordable. The small, fast tiers your kid’s school can pay for exist partly because of it.
  • The teaching frame that works: learning from a teacher makes you able to do new things; copying a teacher’s answers makes you able to reproduce old ones. Distillation is somewhere in between, and where exactly is the interesting question.

What distillation actually does

Start with the original idea. Geoffrey Hinton, Oriol Vinyals, and Jeff Dean’s 2015 paper “Distilling the Knowledge in a Neural Network” describes a compression technique: take an ensemble of large neural networks that performs well but is expensive to run, and transfer what it knows into a single smaller model that’s practical to deploy. Their claim, stated plainly: “it is possible to compress the knowledge in an ensemble into a single model which is much easier to deploy.” They validated it on MNIST and showed significant improvements to acoustic models in a commercial system.

The mechanism is more interesting than “copy the answers.” When a classifier looks at a photo of a dog, it doesn’t just output “dog.” It outputs a probability for every category: 0.90 dog, 0.06 wolf, 0.03 cat, 0.001 truck. That full distribution contains information the single label doesn’t. It says a dog looks somewhat like a wolf and almost nothing like a truck.

Those probabilities are the soft targets, and training a student on them transfers more than training on hard labels does. The student learns the teacher’s sense of what resembles what, not just its final answers.

For language models the same idea applies to next-token predictions, and in practice a lot of modern distillation is simpler: generate a large number of outputs from a strong model, then fine-tune a smaller model on those outputs. Cheaper, cruder, and effective.

Why labs treat it as a threat category

Here’s the commercial reality. Training a frontier model requires enormous compute, data work, and safety evaluation. The resulting model is the product. If a competitor can query it at retail API prices, collect a few million high-quality outputs, and fine-tune a smaller model on them, they’ve acquired a large share of that capability for a rounding error of the original cost.

That’s the concern, and it’s why Anthropic’s June 9, 2026 Fable 5 and Mythos 5 announcement names distillation alongside cybersecurity and bio/chem. The stated mechanism: “When Fable’s classifiers detect a request related to cybersecurity, biology and chemistry, or distillation, the response is automatically handled by Claude Opus 4.8 instead.”

Two details make that design worth understanding. First, it’s a fallback rather than a refusal: the request still gets answered, by a different model, and the user is told. Second, the reported frequency is low, with more than 95% of Fable sessions involving no fallback at all. That’s a narrow filter, not a blanket restriction. Our explainer on how safety classifiers work covers the mechanism.

The same announcement describes Mythos 5 as the same underlying model with safeguards lifted in those areas, restricted to authorized users and initially deployed through Project Glasswing with the US government.

Terms of service at most major labs also prohibit using outputs to train competing models. Whether that’s enforceable in every jurisdiction is a genuinely unsettled legal question, and I’d rather flag the uncertainty than pretend otherwise.

Worth stating clearly: distillation is not illegal or unethical in itself. It’s a standard, published, widely used technique. What labs object to is specifically using their model as an unpaid teacher for a competing product.

Why distillation is also the reason AI is affordable

Now the other side, which is the part that matters most to families.

Every cheap, fast model tier you can actually afford exists partly because of distillation and related compression techniques. The pattern across the industry is a large expensive model plus smaller, cheaper variants that capture much of its capability. Look at the published pricing: Claude Haiku 4.5 at $1 per million input tokens against Claude Fable 5.1 at $10. A tenfold difference in price, and the cheap tier is genuinely useful for most homework.

Three concrete consequences:

Schools can afford AI. A district weighing a tool for tens of thousands of students is choosing between tiers, and the affordable tier is the one that makes the deployment possible at all.

On-device AI exists. A model small enough to run on a phone without an internet connection got small through compression. Google’s August 2026 roundup describes Gemini Nano running on the Pixel 11’s Tensor G6 chip.

Open models get better. Google’s Gemma family passed one billion downloads and runs in environments “from phones and edge infrastructure to space,” per the same roundup. What you can legally do with those weights is set by the Gemma Terms of Use, which we cover in our model licenses explainer. Compression is what makes small open models worth downloading.

So the honest framing for a kid is not “distillation is stealing.” It’s that the same technique is foundational infrastructure when a lab applies it to its own model and a competitive threat when someone applies it to a rival’s. Our piece on why AI models come in different sizes covers the tier landscape.

How to Teach Your Kid About Model Distillation

This maps onto something every kid already has opinions about: copying homework versus learning from a good explanation.

Ages 5–8: The confident guess game

Show your kid a blurry photo or a partially covered picture of an animal. Ask not just “what is it?” but “how sure are you, and what else could it be?” They’ll say something like “a dog, but maybe a wolf, definitely not a bird.” Point out that the second answer taught you more. Then say: “Big computer helpers say answers that way too, with all their maybes, and small helpers learn faster when the big one shares the maybes instead of just the answer.” That’s soft targets, explained to a six-year-old.

Ages 9–12: Teach a sibling two ways

Have your kid teach a younger sibling or a friend something they know, twice, two different ways. First: give them the answers to ten questions and have them memorize. Second: explain why the answers are what they are, including which wrong answers are close and which are absurd. Then quiz the learner on a new, unseen question. The second method transfers; the first doesn’t. Now name it: “The second way is what makes distillation work, and it’s also why copying homework doesn’t help you on the test.”

Ages 13+: Read the classifier decision

Have your teen read Anthropic’s Fable 5 announcement and answer three questions in writing: What are the three categories that trigger a fallback? Why would a company put distillation in the same list as bio/chem and cybersecurity? What does it tell you that the fallback is to a different model rather than a refusal? There’s no single right answer to the second one, and the argument is the exercise. Bonus: have them find the reported percentage of sessions with no fallback and consider what that number is meant to communicate.

The question to ask: “What’s the difference between learning from a teacher and copying a teacher’s answers? Can a computer tell?”

Teacher to student: what transfers and what doesn’t

AspectTeacher modelStudent modelWhat transfers
SizeVery large; expensive to runMuch smaller; cheap and fastNot the size, obviously
Training costEnormous: compute, data, safety evaluationA fraction of the teacher’sOnly the behavior, not the investment
What it learned fromRaw data at scaleThe teacher’s outputs, including probability distributionsHinton et al.’s “soft targets” carry similarity structure
Capability ceilingSets the ceilingGenerally at or below the teacherStudents rarely exceed teachers on the distilled tasks
Novel situationsHandles unfamiliar inputs from broad trainingWeaker where the teacher’s outputs didn’t cover itCoverage gaps become student gaps
Legitimate useThe lab distills its own model into cheap tiersFlash/Haiku/Nano-class productsEveryone benefits: cheaper AI
Contested useA competitor queries it to build a rivalA cheaper clone of someone else’s workThis is what classifiers and terms of service target

The row that matters for a kid’s understanding is the fifth one. A distilled student inherits gaps. It’s good where the teacher’s outputs covered the territory and weaker in the corners nobody sampled, which is exactly what happens to a student who studies from an answer key instead of the material.

What to do at home

Use the copying-versus-learning frame

The distinction between memorizing answers and understanding why answers are right is the single most useful academic concept a kid can hold, and distillation gives you a fresh way to talk about it that isn’t a lecture about cheating.

Point out which tier your kid is using

Cheap tiers are often distilled or compressed versions of bigger models. Knowing that explains why the fast one sometimes misses something the slow one catches, and gives your kid a real reason to switch tiers deliberately.

Ask what happens when the student hits a gap

A distilled model is weaker outside the territory its teacher’s outputs covered. When a cheap model confidently fumbles an unusual question, that’s often a coverage gap, and recognizing it is more useful than concluding “AI is dumb.”

Note who’s on the list

When a company publishes what its classifiers watch for, that list tells you what it considers a real risk. Cybersecurity, bio/chem, and distillation is a revealing set: two safety concerns and one commercial one, treated with the same machinery. That’s worth a teenager’s attention.

What not to do

Don’t teach that distillation is cheating. It’s a published technique from Hinton, Vinyals, and Dean, it’s used by every major lab on its own models, and it’s the reason a school district can afford an AI tool at all. The contested version is narrow: using a competitor’s model as an unpaid teacher. Collapsing the two teaches a kid something false about how the field works.

What to Watch For Over the Next 3 Months

  • Week 4: Your kid can explain teacher and student models in one sentence each, and knows that the cheap tier they use is probably a compressed version of something bigger.
  • Month 2 red flags: They think distillation means the small model is just a copy with nothing lost. Or they conclude that all AI companies are stealing from each other.
  • Month 3 self-check: Ask why a company would treat distillation as seriously as cybersecurity. A reasoned answer about training cost and competition means they’ve got it.

Frequently Asked Questions

What is model distillation in simple terms?

Training a small model to imitate a large one, so most of the capability transfers into something much cheaper to run. Hinton, Vinyals, and Dean introduced the modern version in 2015, showing that the knowledge in an ensemble of large networks can be compressed into a single deployable model.

Why does Claude route distillation requests to a different model?

Per Anthropic’s June 9, 2026 announcement, when Fable 5’s classifiers detect a request related to cybersecurity, biology and chemistry, or distillation, the response is handled by a more restricted model instead, and the user is told. The company reports more than 95% of Fable sessions involve no fallback.

Is distillation illegal?

The technique itself is a published, standard method used across the industry. What’s contested is using a specific provider’s outputs to train a competing model, which most terms of service prohibit. Whether those terms are enforceable everywhere is an unsettled legal question.

Does the small model end up as good as the big one?

Usually close on the tasks it was distilled for, and weaker elsewhere. A student inherits its teacher’s coverage: strong where the teacher’s outputs were sampled, weaker in the corners nobody covered. That’s why cheap tiers can feel uneven on unusual questions.

Should my kid care about any of this?

The mechanism is a clean lesson in what transfers when you learn from someone versus copy them. And practically, knowing that cheap tiers are compressed versions of bigger models explains their uneven behavior better than “it’s the bad one.”


About the author

Ricky Flores is the founder of HiWave Makers and an electrical engineer with 15+ years of experience building consumer technology at Apple, Samsung, and Texas Instruments. He writes about how kids learn to build, think, and create in a tech-saturated world. Read more at hiwavemakers.com.


Sources

  1. Hinton, G., Vinyals, O., & Dean, J. (2015). “Distilling the Knowledge in a Neural Network.” arXiv:1503.02531. https://arxiv.org/abs/1503.02531
  2. Anthropic. (2026, June 9). “Claude Fable 5 and Claude Mythos 5.” https://www.anthropic.com/news/claude-fable-5-mythos-5
  3. Anthropic. (2026). “Pricing.” Claude Platform Documentation. https://platform.claude.com/docs/en/about-claude/pricing
  4. Google. (2026, August). “Google AI updates, August 2026.” The Keyword. https://blog.google/innovation-and-ai/technology/google-ai-updates-august-2026/
  5. Shazeer, N., Mirhoseini, A., Maziarz, K., et al. (2017). “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.” arXiv:1701.06538. https://arxiv.org/abs/1701.06538
  6. MarkTechPost. (2026, September 1). “Anthropic releases Claude Fable 5.1 and Claude Mythos 5.1.” https://www.marktechpost.com/2026/09/01/anthropic-releases-claude-fable-5-1-and-claude-mythos-5-1-52-6-on-terminal-bench-science-and-75-cheaper-cache-reads/
  7. Google AI. “Gemma Terms of Use.” https://ai.google.dev/gemma/terms
Ricky Flores
Written by Ricky Flores

Founder of HiWave Makers and electrical engineer with 15+ years working on projects with Apple, Samsung, Texas Instruments, and other Fortune 500 companies. He writes about how kids learn to build, think, and create in a tech-driven world.