arrow
Return

Spot-Adaptive Knowledge Distillation

delete2022-01-01
delete55
delete
OA
AI
宋杰 (Jie Song)
Y
Ying Chen
J
Jingwen Ye
宋明黎 (Mingli Song) *
DOI:10.1109/TIP.2022.3170728delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Knowledge distillation (KD) has become a well established paradigm for compressing deep neural networks. The typical way of conducting knowledge distillation is to train the student network under the supervision of the teacher network to harness the knowledge at one or multiple spots (i.e., layers) in the teacher network. The distillation spots, once specified, will not change for all the training samples, throughout the whole distillation process. In this work, we argue that distillation spots should be adaptive to training samples and distillation epochs. We thus propose a new distillation strategy, termed spot-adaptive KD (SAKD), to adaptively determine the distillation spots in the teacher network per sample, at every training iteration during the whole distillation period. As SAKD actually focuses on where to distill instead of what to distill that is widely investigated by most existing works, it can be seamlessly integrated into existing distillation methods to further improve their performance. Extensive experiments with 10 state-of-the-art distillers are conducted to demonstrate the effectiveness of SAKD for improving their distillation performance, under both homogeneous and heterogeneous distillation settings. Code is available at https://github.com/zju-vipa/spot-adaptive-pytorch.
Keywords:
Knowledge engineering
Training
Routing
Data models
Adaptation models
Deep learning
Training data
Knowledge distillation
deep neural networks
distillation spots
spot-adaptive distillation

Journal

IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
IF:
13.7
Papers:
1.0W
Citations:
8.4W

Organization

Z
zhejiang university
Scholars:
17.4W
Papers: 12.0W
Citations: 152