arrow
Return

Adaptive teaching with shared classifier for knowledge distillation

delete2025-08-27
delete0
PRE
AI
J
Jaeyeon Jang *
Y
Young-Ik Kim
H
Hyeonseong Lee
DOI:10.1016/j.neucom.2025.131343delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• A novel hybrid knowledge distillation (KD) method is proposed. • The pretrained teacher model self-adjusts its parameters to enhance the KD process. • With the support of a shared classifier, the performance degradation is minimized. • State-of-the-art performance is achieved in both single- and multi-teacher setups.
Keywords:
hybrid knowledge distillation
teacher model self-adjustment
shared classifier
performance degradation minimization
multi-teacher setup

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

T
the catholic university of korea
Scholars:
1.4K
Papers: 569
Citations: 0
A
ai lab
Scholars:
39
Papers: 14
Citations: 0