arrow
返回

Parameter-Efficient and Student-Friendly Knowledge Distillation

delete2024-01-01
delete12
delete
OA
AI
J
Jun Rao
X
Xv Meng
L
Liang Ding
S
Shuhan Qi *
刘
刘学博 (Xuebo Liu)
张
张民 (Min Zhang)
D
Dacheng Tao
DOI:10.1109/TMM.2023.3321480delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Pre-trained models are frequently employed in multimodal learning. However, these models have too many parameters and need too much effort to fine-tune the downstream tasks. Knowledge distillation (KD) is a method to transfer knowledge using the soft label from this pre-trained teacher model to a smaller student, where the parameters of the teacher are fixed (or partially) during training. Recent studies show that this mode may cause difficulties in knowledge transfer due to the mismatched model capacities. To alleviate the mismatch problem, adjustment of temperature parameters, label smoothing and teacher-student joint training methods (online distillation) to smooth the soft label of a teacher network, have been proposed. But those methods rarely explain the effect of smoothed soft labels to enhance the KD performance. The main contributions of our work are the discovery, analysis, and validation of the effect of the smoothed soft label and a less time-consuming and adaptive transfer of the pre-trained teacher's knowledge method, namely PESF-KD by adaptive tuning soft labels of the teacher network. Technically, we first mathematically formulate the mismatch as the sharpness gap between teacher's and student's predictive distributions, where we show such a gap can be narrowed with the appropriate smoothness of the soft label. Then, we introduce an adapter module for the teacher and only update the adapter to obtain soft labels with appropriate smoothness. Experiments on various benchmarks including CV and NLP show that PESF-KD can significantly reduce the training cost while obtaining competitive results compared to advanced online distillation methods.
Keyword:
Training
Smoothing methods
Knowledge transfer
Data models
Adaptation models
Predictive models
Knowledge engineering
Knowledge distillation
parameter-efficient
image classification

期刊

IEEE Transactions on Multimedia 封面图
IEEE Transactions on Multimedia
IF:
9.7
论文数:
4.5K
被引数:
2.4W

机构

H
harbin institute of technology
学者数:
8.0W
论文数: 6.6W
被引数: 66
U
University of Sydney
学者数:
6.5W
论文数: 6.2W
被引数: 90
引用论文

引用论文

Dihydrodiborolyl-transition metal-carboranyl triple-decker sandwich complexes
err2002-05-01
err0
PREAI
errMartin D. Attwood; Kathleen K. Fonda; Russell N. Grimes; Gregor Brodt; Dongqi Hu; Ulrich Zenneck; Walter Siebert
err分享
err收藏
Knowledge Distillation: A Survey知识蒸馏: 一项调查
err2021-03-22
err1.5K
PREAI
errGou, Jianping; Yu, Baosheng; Maybank, Stephen J.; Tao, Dacheng
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容