A model-training technique where a small 'student' model learns to mimic the outputs of a large 'teacher' model — producing a much smaller model that retains most of the teacher's quality.
Knowledge Distillation trains a student model on the teacher's output distributions (soft targets) rather than just the ground-truth labels (hard targets). The soft targets carry richer information — the teacher's confidence across all classes, not just the correct one. Used to compress LLMs (Llama → distilled smaller variants), to fine-tune cheap models on outputs from expensive ones (GPT-4 → custom smaller models), and in computer vision compression.
Distilling GPT-4 responses into a fine-tuned Llama 3 8B model — getting 80% of GPT-4 quality at 1/30th the inference cost.
Distillation is how teams ship LLM-quality results at production-friendly cost — train once with the big model, infer cheaply with the small one.
Need help implementing this in your business?
Get Started