arrow
返回

Learning Low Resource Consumption CNN Through Pruning and Quantization

delete2021-01-01
delete9
PRE
AI
戚琦 (Qi Qi)
Y
Yan Lu
J
Jiashi Li
王晶钰 (Jingyu Wang) *
H
Haifeng Sun
J
Jianxin Liao
DOI:10.1109/TETC.2021.3050770delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Deep learning models have evolved into powerful tools that can be used for many artificial intelligence tasks. However, deploying deep neural networks into real-world applications is still challenging due to their high computational complexity and storage overhead. Fortunately, a densely connected neural network can be converted into a sparsely connected network with low resource demand by the neural network compression. Since deep neural networks are complicated, compression mechanism should find a tradeoff between compression ratio and model accuracy. In this article, by analyzing the statistics of channel connection, we propose an interactive neural network compression mechanism including out-in-channel pruning and neural network quantization. Many channel pruning works apply structured sparsity regularization on each layer separately. We consider correlations between successive layers to retain predictive power of the compact network. A global greedy pruning algorithm is designed to remove redundant out-in-channels in an iterative way. Moreover, in order to solve the shortcomings of the one-shot quantization, we propose the incremental quantization algorithm in the dimension of the output channel, which can smooth network fluctuations and recover accuracy better during retraining. Our mechanism is comprehensively evaluated with various Convolutional Neural Networks (CNN) architectures on popular datasets. Notably, on ImageNet-1K, the out-in-channel pruning reduce 54.0 percent FLOPS on AlexNet and 50.0 percent FLOPs on ResNet-50 with only 0.15 and 0.37 percent top-1 accuracy drop respectively. On classification and style transfer tasks, the superiority of incremental quantization increases with the decrease of the number of quantization bits.
Keyword:
Neural networks
Quantization (signal)
Computational modeling
Training
Task analysis
Deep learning
Redundancy
Deep neural networks
channel pruning
network quantization
low resource consumption
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Emerging Topics in Computing 封面图
IEEE Transactions on Emerging Topics in Computing
IF:
5.4
论文数:
1.1K
被引数:
3.4K

机构

B
beijing university of posts & telecommunications
学者数:
1.4W
论文数: 1.2W
被引数: 9
引用论文

引用论文

Calcium-modulating cyclophilin ligand regulates membrane trafficking of postsynaptic GABAA receptors
err2008-06-01
err0
errOAAI
errXu Yuan; Jun Yao; David Norris; David D. Tran; Richard J. Bram; Gong Chen; Bernhard Luscher
err分享
err收藏
Utilización de medicinas alternativas y consumo de drogas por pacientes con enfermedad inflamatoria intestinal
err2007-01-01
err0
PREAI
errEsther García-Planella; Laura Marín; Eugeni Domènech; Isabel Bernal; Míriam Mañosa; Yamile Zabana; Miquel A. Gassull
err分享
err收藏
Major Adverse Limb Events and Mortality in Patients With Peripheral Artery Disease
err2018-05-01
err0
errOAAI
errSonia S. Anand; Francois Caron; John W. Eikelboom; Jackie Bosch; Leanne Dyal; Victor Aboyans; Maria Teresa Abola; Kelley R.H. Branch; Katalin Keltai; Deepak L. Bhatt; Peter Verhamme; Keith A.A. Fox; Nancy Cook-Bruns; Vivian Lanius; Stuart J. Connolly; Salim Yusuf
err分享
err收藏
Industrial Cyber-Physical Systems-Based Cloud IoT Edge for Federated Heterogeneous Distillation
err2021-08-01
err42
errOAAI
errWang, Chengjia; Yang, Guang; Papanastasiou, Giorgos; Zhang, Heye; Rodrigues, Joel J. P. C.; de Albuquerque, Victor Hugo C.
err分享
err收藏
学者 查看更多内容