arrow
Return

Decoupled Classifier Knowledge Distillation

delete2025-02-21
delete0
delete
OA
AI
W
Wang, Hairui
M
Mengjie Dong
G
Guifu Zhu
Y
Ya Li *
DOI:10.1371/journal.pone.0314267delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Mainstream knowledge distillation methods primarily include self-distillation, offline distillation, online distillation, output-based distillation, and feature-based distillation. While each approach has its respective advantages, they are typically employed independently. Simply combining two distillation methods often leads to redundant information. If the information conveyed by both methods is highly similar, this can result in wasted computational resources and increased complexity. To provide a new perspective on distillation research, we aim to explore a compromise solution that aligns complex features without conflicting with output alignment. In this work, we propose to decouple the classifier's output into two components: non-target classes learned by the student, and target classes obtained by both the teacher and the student. Finally, we introduce Decoupled Classifier Knowledge Distillation (DCKD), where on one hand, we fix the correct knowledge that the student has already acquired, which is crucial for merging the two methods; on the other hand, we encourage the student to further align its output with that of the teacher. Compared to using a single method, DCKD achieves superior results on both the CIFAR-100 and ImageNet datasets for image classification and object detection tasks, without reducing training efficiency. Moreover, it allows relational-based and feature-based distillation to operate more efficiently and flexibly. This work demonstrates the great potential of integrating distillation methods, and we hope it will inspire future research.

Journal

PLoS One cover
PLoS One
IF:
2.6
Papers:
2.6W
Citations:
81.6W

Organization

K
kunming univ sci &technol
Scholars:
3.1K
Papers: 1.1K
Citations: 13
Cited Papers

Cited Papers

Knowledge Distillation: A Survey
err2021-03-22
err1.5K
PREAI
errGou, Jianping; Yu, Baosheng; Maybank, Stephen J.; Tao, Dacheng
errShare
errSave
Model compression
err2006-08-20
err0
PREAI
errCristian Buciluǎ; Rich Caruana; Alexandru Niculescu-Mizil
errShare
errSave
Cross-Layer Distillation with Semantic Calibration
err2021-05-18
err0
errOAAI
errDefang Chen; Jian-Ping Mei; Yuan Zhang; Can Wang; Zhe Wang; Yan Feng; Chun Chen
errShare
errSave
Frequency Attention for Knowledge Distillation
err2024-01-03
err0
errOAAI
errCuong Pham; Van-Anh Nguyen; Trung Le; Dinh Phung; Gustavo Carneiro; Thanh-Toan Do
errShare
errSave
ImageNet Large Scale Visual Recognition Challenge
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
errShare
errSave
A Comprehensive Overhaul of Feature Distillation
err2019-10-01
err0
errOAAI
errByeongho Heo; Jeesoo Kim; Sangdoo Yun; Hyojin Park; Nojun Kwak; Jin Young Choi
errShare
errSave
Student-friendly knowledge distillation
err2024-07-01
err5
errOAAI
errYuan, Mengyang; Lang, Bo; Quan, Fengnan
errShare
errSave
Improving Knowledge Distillation via Regularizing Feature Direction and Norm
err2024-11-03
err0
PREAI
errYuzhu Wang; Lechao Cheng; Manni Duan; Yongheng Wang; Zunlei Feng; Shu Kong
errShare
errSave
researcher View more