arrow
Return

SAKD: Sparse attention knowledge distillation

delete2024-06-01
delete2
PRE
AI
Z
Zhen Guo
P
Pengzhou Zhang *
L
Liang Peng
DOI:10.1016/j.imavis.2024.105020delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep learning techniques have gained significant interest due to their success in large model scenarios. However, large models often require massive computational resources, which can challenge end devices with limited storage capabilities. Transferring knowledge from big to small models and achieving similar results with limited resources requires further research. Knowledge distillation techniques, which involve using teacher-student models to migrate large model capabilities to small models, have been widely used in model compression and knowledge transfer. In this paper, a novel knowledge distillation approach is proposed, which utilizes the sparse attention mechanism (SAKD). SAKD computes attention using student features as queries and teacher features as key values and performs sparse attention values by random deactivation. Then, this sparse attention value is used to reweight the feature distance of each teacher-student feature pair to avoid negative transfer. Comprehensive experiments demonstrate the effectiveness and generality of our approach. Moreover, our SAKD method outperforms previous state-of-the-art methods on image classification tasks.
Keywords:
Knowledge distillation
Attention mechanisms
Sparse attention mechanisms

Journal

Image and Vision Computing cover
Image and Vision Computing
IF:
4.2
Papers:
4.0K
Citations:
6.7K

Organization

C
Communication University of China
Scholars:
1.1K
Papers: 820
Citations: 326