arrow
返回

Multi-target Knowledge Distillation via Student Self-reflection

delete2023-04-25
delete24
delete
OA
AI
J
Jianping Gou
X
Xiangshuo Xiong
B
Baosheng Yu *
L
Lan Du
Y
Yibing Zhan
D
Dacheng Tao
DOI:10.1007/s11263-023-01792-zdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Knowledge distillation is a simple yet effective technique for deep model compression, which aims to transfer the knowledge learned by a large teacher model to a small student model. To mimic how the teacher teaches the student, existing knowledge distillation methods mainly adapt an unidirectional knowledge transfer, where the knowledge extracted from different intermedicate layers of the teacher model is used to guide the student model. However, it turns out that the students can learn more effectively through multi-stage learning with a self-reflection in the real-world education scenario, which is nevertheless ignored by current knowledge distillation methods. Inspired by this, we devise a new knowledge distillation framework entitled multi-target knowledge distillation via student self-reflection or MTKD-SSR, which can not only enhance the teacher's ability in unfolding the knowledge to be distilled, but also improve the student's capacity of digesting the knowledge. Specifically, the proposed framework consists of three target knowledge distillation mechanisms: a stage-wise channel distillation (SCD), a stage-wise response distillation (SRD), and a cross-stage review distillation (CRD), where SCD and SRD transfer feature-based knowledge (i.e., channel features) and response-based knowledge (i.e., logits) at different stages, respectively; and CRD encourages the student model to conduct self-reflective learning after each stage by a self-distillation of the response-based knowledge. Experimental results on five popular visual recognition datasets, CIFAR-100, Market-1501, CUB200-2011, ImageNet, and Pascal VOC, demonstrate that the proposed framework significantly outperforms recent state-of-the-art knowledge distillation methods.
Keyword:
Knowledge distillation
Self-reflection learning
Model compression
Deep learning

期刊

International Journal of Computer Vision 封面图
International Journal of Computer Vision
IF:
9.3
论文数:
3.9K
被引数:
2.8W

机构

J
Jiangsu University
学者数:
4.0W
论文数: 2.8W
被引数: 5.5W
M
Monash University
学者数:
5.4W
论文数: 5.4W
被引数: 79
S
southwest university - china
学者数:
2.6W
论文数: 1.9W
被引数: 21
U
University of Sydney
学者数:
6.5W
论文数: 6.2W
被引数: 90
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
err分享
err收藏
err分享
err收藏
Knowledge Distillation: A Survey知识蒸馏: 一项调查
err2021-03-22
err1.5K
PREAI
errGou, Jianping; Yu, Baosheng; Maybank, Stephen J.; Tao, Dacheng
err分享
err收藏
Collaborative Knowledge Distillation via Multiknowledge Transfer
err2024-05-01
err11
PREAI
errGou, Jianping; Sun, Liyuan; Yu, Baosheng; Du, Lan; Ramamohanarao, Kotagiri; Tao, Dacheng
err分享
err收藏
A role for anion transport in the regulation of release from chromaffin granules and exocytosis from cells
err2004-02-19
err0
PREAI
errHarvey B. Pollard; Christopher J. Pazoles; Carl E. Creutz; Avner Ramu; Charles A. Strott; Probhati Ray; Edward M. Brown; G. D. Aurbach; Karen M. Tack‐Goldman; N. Raphael Shulman
err分享
err收藏
学者 查看更多内容