How Big Models Teach Small Models, and Why It Looks Like School
Knowledge distillation is a large model teaching a small one, and it works almost exactly like school: soft labels instead of a bare answer key, pairing weeks instead of reports, and a tutor who marks your own attempt rather than a textbook of last year’s solutions. The limits map too, and the last one costs money when you size the hardware.





