返回
SCL-IKD: intermediate knowledge distillation via supervised contrastive representation learning
DOI:10.1007/s10489-023-05036-y.png)
摘要
En 中文
Knowledge distillation, which extracts dark knowledge from a deep teacher model to drive the learning of a shallow student model, is helpful in several tasks, including model compression and regularization. While previous research has focused on architecture-driven solutions for extracting information from the teacher models, these solutions are focused on a single task and fail to extract rich dark knowledge from large teacher networks in the presence of capacity gaps for broader applications. Hence, in this paper, we propose a supervised contrastive learning-based intermediate knowledge distillation (SCL-IKD) technique that is more effective in distilling knowledge from teacher networks to train a student model for classification tasks. SCL-IKD, unlike other approaches, is model agnostic and may be used in a variety of teacher-student cross-architectures. Investigations on several datasets reveal that SCL-IKD can achieve 3-4%\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$3-4\%$$\end{document} better top-1 accuracy over several state-of-the-art baselines. Furthermore, compared to the baselines, SCL-IKD is found better to handle capacity gaps between teacher and student models and is significantly more robust to symmetric noisy labels and data availability.
Keyword:
Knowledge distillation
Image classification
Supervised contrastive learning
Dark knowledge extraction
Deep learning
期刊
IF:
3.5
论文数:
7.6K
被引数:
1.7W

