Distillation is a way to shrink AI down to size. A large, powerful model acts as a teacher, and a much smaller student model learns to copy its answers. The result is a compact model that captures a lot of the teacher's ability while running far faster and cheaper, which is handy for phones and everyday apps.
For example, A giant flagship model can be distilled into a lightweight version that runs smoothly on a laptop.