arrow
Return

Teacher-student complementary sample contrastive distillation

delete2024-02-01
delete2
PRE
AI
Z
Zhiqiang Bao
Z
Zhenhua Huang *
J
Jianping Gou
L
Lan Du
K
Kang Liu
J
Jingtao Zhou
Y
Yunwen Chen
DOI:10.1016/j.neunet.2023.11.036delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Knowledge distillation (KD) is a widely adopted model compression technique for improving the performance of compact student models, by utilizing the dark knowledgeof a large teacher model. However, previous studies have not adequately investigated the effectiveness of supervision from the teacher model, and overconfident predictions in the student model may degrade its performance. In this work, we propose a novel framework, Teacher-Student Complementary Sample Contrastive Distillation (TSCSCD), that alleviate these challenges. TSCSCD consists of three key components: Contrastive Sample Hardness (CSH), Supervision Signal Correction (SSC), and Student Self-Learning (SSL). Specifically, CSH evaluates the teacher's supervision for each sample by comparing the predictions of two compact models, one distilled from the teacher and the other trained from scratch. SSC corrects weak supervision according to CSH, while SSL employs integrated learning among multi-classifiers to regularize overconfident predictions. Extensive experiments on four real-world datasets demonstrate that TSCSCD outperforms recent state-of-the-art knowledge distillation techniques.
Keywords:
Knowledge distillation
Transfer learning
Model regularization
Sample hardness
Deep learning

Journal

Neural Networks cover
Neural Networks
IF:
6.3
Papers:
7.8K
Citations:
3.0W

Organization

M
Monash University
Scholars:
5.4W
Papers: 5.4W
Citations: 79
S
southwest university - china
Scholars:
2.6W
Papers: 1.9W
Citations: 21
S
south china normal university
Scholars:
2.0W
Papers: 1.3W
Citations: 13
researcher View more organizations