October 4, 2026
Scaling Inference

Distillation for Production: Shrinking Frontier Models Without Losing Accuracy

Distillation for Production: Shrinking Frontier Models Without Losing Accuracy

Distillation trains a small student model to reproduce a large teacher’s behaviour, not merely its final answers. The gain comes from soft targets: the teacher’s full probability distribution carries far more information than a hard label. Narrow the task enough and a distilled student can match its teacher at a fraction of the serving cost.

Consider a classification service running a frontier model on ten million documents a month. It works, and the bill is the largest line in the team’s budget.

The task itself is narrow. The model is being paid for general capability it never uses. That gap between what you buy and what you need is exactly what distillation closes.

At a glance

  • Soft targets carry more signal than hard labels, which is the core mechanism
  • Distillation works best on narrow tasks and worst on open-ended ones
  • Coverage of the input distribution matters more than volume of examples
  • A student inherits the teacher’s errors along with its skill
  • Check licence terms before training on another model’s outputs
  • Distilled models fail hardest on inputs the teacher never demonstrated

Why do soft targets matter so much?

Because they encode how the teacher was wrong, not only what it decided.

Hinton, Vinyals and Dean framed this as the dark knowledge in a model’s output distribution (arXiv:1503.02531). A hard label says an image is a dog. A soft distribution says it is mostly dog, slightly wolf, marginally fox, and essentially never truck.

That relative structure teaches the student which categories neighbour each other. Learning from hard labels alone discards it.

Figure 1 — The same example, two supervision signals

Hard label

The correct class, and nothing else. Every incorrect class is equally wrong.

The student learns the decision but none of the structure behind it, so it must rediscover category relationships from scratch.

One bit of structure

Soft target

The teacher’s full distribution across classes, including the near-misses it considered.

The student inherits a similarity map for free, which is why distillation reaches accuracy that direct training on the same data does not.

Rich structure

This is why distillation beats simply training a small model on the same dataset. The teacher is not supplying answers; it is supplying a shaped view of the problem space that no label set contains.

Where does distillation actually work?

Task shapeSuitabilityWhy
Fixed-label classificationExcellentNarrow output space, dense supervision signal
Structured extractionExcellentWell-defined target, easy to verify
Routing and triageExcellentSmall decision space, high volume
Summarisation in a fixed formatGoodConstrained enough to learn reliably
Domain question answeringModerateWorks if the domain is genuinely bounded
Open-ended reasoningPoorThe student lacks the capacity the task requires
General assistancePoorNo bounded distribution to cover

The pattern is that distillation transfers narrow competence well and broad capability badly. Teams that attempt to distil a general assistant into a small model are trying to compress something whose value is precisely its breadth.

What does the pipeline look like?

  1. Define the task narrowly and write the evaluation before anything else. A vague task cannot be distilled because you cannot tell when it worked.
  2. Collect real input distribution. Production traffic, not invented examples. Coverage of the inputs you actually see matters more than volume.
  3. Generate teacher outputs across that distribution, capturing full probability distributions where the API exposes them.
  4. Include the hard cases deliberately. Edge cases and ambiguous inputs are where the student will fail, so they must be represented.
  5. Train the student on teacher outputs, using soft targets where available and generated text where not.
  6. Evaluate against the teacher, not against a benchmark. The teacher is the bar you are trying to reach.
  7. Measure the gap on slices, especially rare inputs, since aggregate agreement hides exactly where students break.

Watch for this

Check the terms of service before training a student on another provider’s outputs. Several major providers explicitly prohibit using their model’s output to train competing models, and the restriction is often overlooked because the technical path is so straightforward. This is a legal question that should be settled before the engineering work starts, not after.

What does a distilled student lose?

Three things, all invisible on in-distribution evaluation.

  • Out-of-distribution behaviour. On inputs unlike anything in the training set, the student has nothing to fall back on where the teacher had general capability.
  • Graceful degradation. A large model handling an unfamiliar case tends to produce something reasonable. A small distilled model tends to produce something confidently wrong.
  • Inherited errors. The student learns the teacher’s mistakes as faithfully as its successes, and cannot detect that it has done so.

The third point deserves emphasis. Distillation is a copying process, so teacher bias, teacher blind spots and teacher hallucination patterns all transfer. Validating the student against the teacher will never reveal this, because they agree.

That is a case for at least one evaluation slice with ground truth established independently of the teacher.

How is this different from fine-tuning?

They overlap, and the distinction is about the supervision source.

Fine-tuning adapts a model using human-authored examples. Distillation adapts it using another model’s behaviour. In practice teams often do both: distil to transfer the task, then fine-tune on human-corrected cases where the student diverges.

The relationship to model collapse is worth naming here, since distillation is a single hop from a teacher trained on real data rather than an iterated loop. Repeated distillation across several generations, however, starts to resemble the recursive pattern that degrades distributions, so treat generation depth as something to track.

Where do teams go wrong?

Scoping the task too broadly

The most common failure, and it is decided before any training runs. A student asked to replace a frontier model across five loosely-related tasks will underperform on all five. Distil one task well, then repeat.

Training on convenient rather than representative data

Teacher outputs generated from a tidy sample produce a student that works on tidy inputs. Production traffic is not tidy, and the gap appears immediately after launch.

Declaring success on aggregate agreement

Ninety-five percent agreement with the teacher sounds excellent and can hide total failure on the five percent that matters most. Report agreement per input slice, weighted by how consequential each slice is.

What do experienced teams do differently?

They keep the teacher in production as a fallback rather than switching over entirely.

A confidence threshold on the student, escalating uncertain cases to the teacher, captures most of the cost saving while preserving quality on hard inputs. The economics work because hard inputs are usually a small fraction of traffic.

They also re-distil on a schedule. As the input distribution drifts, a student trained on last year’s traffic degrades in ways that are gradual and hard to notice. Treating distillation as a recurring pipeline rather than a one-off project is what keeps it working.

A short glossary

Teacher and student
The large source model whose behaviour is copied, and the smaller model being trained to reproduce it.
Soft targets
The teacher’s full output probability distribution, used as a training signal in place of hard labels.
Dark knowledge
The relative structure encoded in a model’s non-selected probabilities, revealing which classes it considers similar.
Temperature
A scaling factor applied to logits during distillation to expose more of the distribution’s structure.
Confidence routing
Serving most traffic from the student and escalating low-confidence cases to the teacher.

Key takeaways

  • Distillation transfers behaviour through soft targets, which carry more information than hard labels.
  • Hinton, Vinyals and Dean described this as dark knowledge in the model’s output distribution.
  • Narrow, well-defined tasks distil well; open-ended reasoning and general assistance do not.
  • Coverage of the real input distribution matters more than the number of training examples.
  • Students inherit teacher errors and cannot detect them, so validate against independent ground truth on at least one slice.
  • Check provider terms before training on another model’s output, since many prohibit it.
  • Keep the teacher as a confidence-triggered fallback rather than switching over completely.

Frequently asked questions

What makes distillation better than training a small model directly?

Soft targets. The teacher’s full probability distribution shows which categories it considered and how close they were, which encodes structure no hard label set contains. The student inherits that similarity map, reaching accuracy that direct training on the same data typically does not.

Which tasks distil well?

Narrow ones with bounded output spaces: fixed-label classification, structured extraction, routing and triage, and format-constrained summarisation. Open-ended reasoning and general assistance distil badly, because their value lies precisely in the breadth a small student cannot hold.

How much data do I need to distil a model?

Coverage matters more than volume. A modest dataset spanning your real input distribution, including edge cases and ambiguous inputs, outperforms a much larger set of easy examples. Generate teacher outputs from production traffic rather than from a convenient tidy sample.

Does the student inherit the teacher’s mistakes?

Yes, faithfully, and it cannot tell that it has. Teacher bias, blind spots and hallucination patterns all transfer. Validating the student against the teacher will never surface this because they agree, so keep at least one evaluation slice with independently established ground truth.

Is distillation the same as fine-tuning?

They overlap but differ in supervision source. Fine-tuning uses human-authored examples; distillation uses another model’s behaviour. Many teams do both, distilling to transfer the task and then fine-tuning on human-corrected cases where the student diverges from what they want.

Is it legal to distil from a commercial API?

Often not, and this needs checking first. Several major providers explicitly prohibit using their outputs to train competing models. The technical path is easy enough that teams frequently build before reading the terms, which is an expensive order to do things in.

Should I replace the teacher entirely once the student works?

Usually not. Serve most traffic from the student and escalate low-confidence cases to the teacher. Since hard inputs are typically a small share of volume, this captures most of the cost saving while preserving quality exactly where the student is weakest.

Does repeated distillation cause model collapse?

A single hop from a teacher trained on real data is low risk. Repeated distillation across several generations begins to resemble the recursive training pattern that narrows distributions and erases rare cases, so track how many generations removed from real data your student actually is.

References

    Avatar photo
    Sophie Williams first earned a First-Class Honours degree in Electrical Engineering from the University of Manchester, then a Master's degree in Artificial Intelligence from the Massachusetts Institute of Technology (MIT). Over the past ten years, Sophie has become quite skilled at the nexus of artificial intelligence research and practical application. Starting her career in a leading Boston artificial intelligence lab, she helped to develop projects including natural language processing and computer vision.From research to business, Sophie has worked with several tech behemoths and creative startups, leading AI-driven product development teams targeted on creating intelligent solutions that improve user experience and business outcomes. Emphasizing openness, fairness, and inclusiveness, her passion is in looking at how artificial intelligence might be ethically included into shared technologies.Regular tech writer and speaker Sophie is quite adept in distilling challenging AI concepts for application. She routinely publishes whitepapers, in-depth pieces for well-known technology conferences and publications all around, opinion pieces on artificial intelligence developments, ethical tech, and future trends. Sophie is also committed to supporting diversity in tech by means of mentoring programs and speaking events meant to inspire the next generation of female engineers.Apart from her job, Sophie enjoys rock climbing, working on creative coding projects, and touring tech hotspots all around.

      Leave a Reply

      Your email address will not be published. Required fields are marked *