This work reduces the computational cost of large-scale knowledge distillation.
problem Heavy computational costs in training large-scale knowledge distillation models.
method Dynamic Importance Sampling applied to the interaction between teacher and student.
result Our method reduces training time while maintaining competitive performance.