arrow
返回

Inference-aware convolutional neural network pruning

delete2022-10-01
delete16
PRE
AI
T
Tejalal Choudhary
V
Vipul Kumar Mishra *
A
Anurag Goswami
S
S. Jagannathan
DOI:10.1016/j.future.2022.04.031delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Deep neural networks (DNNs) have become an important tool in solving various problems in numerous disciplines. However, DNNs are also known for their high resource requirement, weight redundancy, and large-scale parameters. As a result, the use of DNNs is restricted for devices that lack the necessary resources required to execute, especially resource-constrained devices such as mobile phones, wearable devices, and other edge devices. In recent years, pruning has emerged as an essential technique to reduce insignificant parameters and accelerate the model performance. However, finding the optimal number of parameters that can be pruned without significantly affecting the model performance is a time-consuming, tedious task and require a lot of manual tuning. This paper represent pruning as an optimization problem with the goal of improving DNN run-time inference performance by pruning low impacting parameters (filters) and their corresponding feature maps. To do this, we present a Bayesian optimization-based method for automatically determining the appropriate number of filters for each convolutional layer. Also, we proposed an objective function incorporating distinct model performance and resource-specific constraints. The proposed method is applied to two different kinds of convolutional network architectures (i.e., VGG16 and deeper network ResNet34) on CIFAR10, CIFAR100, and ImageNet datasets. The large-scale ImageNet experimental findings showed that the floating-point operations of the ResNet34 and VGG16 could be reduced by 35.46 percent and 84.97 percent, respectively, with negligible loss of accuracy. (c) 2022 Elsevier B.V. All rights reserved.
Keyword:
Model compression and acceleration
Convolutional neural network
Filter pruning
Bayesian optimization
Resource-constrained devices
Efficient inference

期刊

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
论文数:
6.8K
被引数:
2.3W

机构

University of Missouri System 封面图
University of Missouri System
学者数:
2.9W
论文数: 2.7W
被引数: 75
引用论文

引用论文

Sparse low rank factorization for deep neural network compression
err2020-07-01
err87
PREAI
errSwaminathan, Sridhar; Garg, Deepak; Kannan, Rajkumar; Andres, Frederic
err分享
err收藏
Taking the Human Out of the Loop: A Review of Bayesian Optimization将人类带出循环: 贝叶斯优化的回顾
err2016-01-01
err3.5K
PREAI
errShahriari, Bobak; Swersky, Kevin; Wang, Ziyu; Adams, Ryan P.; de Freitas, Nando
err分享
err收藏
Major Adverse Limb Events and Mortality in Patients With Peripheral Artery Disease
err2018-05-01
err0
errOAAI
errSonia S. Anand; Francois Caron; John W. Eikelboom; Jackie Bosch; Leanne Dyal; Victor Aboyans; Maria Teresa Abola; Kelley R.H. Branch; Katalin Keltai; Deepak L. Bhatt; Peter Verhamme; Keith A.A. Fox; Nancy Cook-Bruns; Vivian Lanius; Stuart J. Connolly; Salim Yusuf
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
学者 查看更多内容