Return
Adaptive teaching with shared classifier for knowledge distillation
DOI:10.1016/j.neucom.2025.131343.png)
Abstract
En 中文
• A novel hybrid knowledge distillation (KD) method is proposed. • The pretrained teacher model self-adjusts its parameters to enhance the KD process. • With the support of a shared classifier, the performance degradation is minimized. • State-of-the-art performance is achieved in both single- and multi-teacher setups.
Keywords:
hybrid knowledge distillation
teacher model self-adjustment
shared classifier
performance degradation minimization
multi-teacher setup
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

