arrow
返回

Multi-perspective analysis on data augmentation in knowledge distillation

delete2024-05-01
delete1
PRE
AI
李
李潍 (Wei Li) *
S
Shitong Shao
Z
Ziming Qiu
A
Aiguo Song
DOI:10.1016/j.neucom.2024.127516delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Knowledge distillation stands as a capable technique for transferring knowledge from a larger to a smaller model, thereby notably enhancing the smaller model's performance. In the recent past, data augmentation has been employed in contrastive learning based knowledge distillation techniques yielding superior results. Despite the significant role of data augmentation, its value remains underappreciated within the domain of knowledge distillation, with no in-depth analysis in the literature thus far. To make up for this oversight, we conduct a multi -perspective theoretical and experimental analysis on the role that data augmentation can play in knowledge distillation. We summarize the properties of data augmentation and list the core findings as follows. (a) Our investigations validate that data augmentation significantly boosts the performance of knowledge distillation on the tasks of image classification and object detection. And this holds true even if the teacher model lacks comprehensive information about the augmented samples. Moreover, our novel J oint D ata A ugmentation (JDA) approach outperforms single data augmentation in knowledge distillation. (b) The pivotal role of data augmentation in knowledge distillation can be theoretically explained via Sharpness -Aware Minimization. (c) The compatibility of data augmentation with various knowledge distillation methods can enhance their performance. In light of these observations, we propose a new method called C osine C onfidence D istillation (CCD) for more reasonable knowledge transfer from augmented samples. Experimental results not only demonstrate that CCD becomes the state-of-the-art method with less storage requirement on CIFAR-100 and ImageNet-1k, but also validate the superiority of CCD over DIST on the object detection benchmark dataset, MS-COCO.
Keyword:
Knowledge distillation
Data augmentation
Decoupling
Sharpness-aware minimization
Few-shot scenarios

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

S
southeast university - china
学者数:
5.3W
论文数: 4.9W
被引数: 57
引用论文

引用论文

Knowledge Distillation: A Survey知识蒸馏: 一项调查
err2021-03-22
err1.5K
PREAI
errGou, Jianping; Yu, Baosheng; Maybank, Stephen J.; Tao, Dacheng
err分享
err收藏
ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
Use and validation of ADMS-Urban in contrasting urban and industrial locations
err2000-01-01
err0
PREAI
errD.J. Carruthers; H.A. Edmunds; A.E. Lester; C.A. McHugh; R.J. Singles
err分享
err收藏
Feature Estimations Based Correlation Distillation for Incremental Image Retrieval
err2022-01-01
err18
errOAAI
errChen, Wei; Liu, Yu; Pu, Nan; Wang, Weiping; Liu, Li; Lew, Michael S.
err分享
err收藏
没有更多内容