arrow
返回

Knowledge Distillation With Feature Self Attention

delete2023-01-01
delete3
delete
OA
AI
S
Sin-Gu Park
D
Dong‐Joong Kang *
DOI:10.1109/ACCESS.2023.3265382delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
With the rapid development of deep learning technology, the size and performance of the network continuously grow, making network compression essential for commercial applications. In this paper, we propose a Feature Self Attention (FSA) module that extracts correlation information between the hidden features of a network and a new method for distilling the correlation features to compress the model. FSA does not require a special module or network to match features between the teacher model and the student model. By removing the multi-head structure and the repeated self-attention blocks in the existing self-attention mechanism, it minimizes the addition of parameters. Based on ResNet-18, 34, the added parameters are only 2.00M and the training speed is also the fastest in comparison to benchmark models. It was demonstrated through experiments that the use of interrelationship loss between features can be beneficial for training student models, indicating the importance of considering correlation information in deep neural network compression. And it was verified through training from scratch on the vanilla without the pre-trained weight of the student model.
Keyword:
Knowledge distillation
self-attention
model compression
training from scratch

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

P
pusan national university
学者数:
2.1W
论文数: 1.9W
被引数: 20
引用论文

引用论文

err分享
err收藏
err分享
err收藏
err分享
err收藏
Atmospheric pressure air plasma treatment to improve the 3D printing of polyoxymethylene
err2019-04-12
err0
PREAI
errIgnacio Muro‐Fraguas; Elisa Sainz‐García; Alpha Pernía‐Espinoza; Fernando Alba‐Elías
err分享
err收藏
err分享
err收藏
学者 查看更多内容