Knowledge Distillation
Training a smaller 'student' model to mimic a larger 'teacher' model, retaining most quality at lower cost.
Many fast, cheap "mini" or "flash" models are distilled from a larger flagship, trading a little accuracy for much lower latency and cost — ideal for high-volume production use.