arrow
返回

Mixed-precision quantized neural networks with progressively decreasing bitwidth

delete2021-03-01
delete27
PRE
AI
T
Tianshu Chu
L
Luo, Qin
J
Jie Yang
X
Xiaolin Huang *
DOI:10.1016/j.patcog.2020.107647delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Efficient model inference is an important and practical issue in the deployment of deep neural networks on resource constraint platforms. Network quantization addresses this problem effectively by leveraging low-bit representation and arithmetic that could be conducted on dedicated embedded systems. In the previous works, the parameter bitwidth is set homogeneously and there is a trade-off between superior performance and aggressive compression. Actually, the stacked network layers, which are generally regarded as hierarchical feature extractors, contribute diversely to the overall performance. For a well-trained neural network, the feature distributions of different categories are organized gradually as the network propagates forward. Hence the capability requirement on the subsequent feature extractors is reduced. It indicates that the neurons in posterior layers could be assigned with lower bitwidth for quantized neural networks. Based on this observation, a simple yet effective mixed-precision quantized neural network with progressively decreasing bitwidth is proposed to improve the trade-off between accuracy and compression. Extensive experiments on typical network architectures and benchmark datasets demonstrate that the proposed method could achieve better or comparable results while reducing the memory space for quantized parameters by more than 25% in comparison with the homogeneous counterparts. In addition, the results also demonstrate that the higher-precision bottom layers could boost the 1-bit network performance appreciably due to a better preservation of the original image information while the lower-precision posterior layers contribute to the regularization of k-bit networks. (C) 2020 Elsevier Ltd. All rights reserved.
Keyword:
Model compression
Quantized neural networks
Mixed-precision
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Pattern Recognition 封面图
Pattern Recognition
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

S
shanghai jiao tong university
学者数:
15.7W
论文数: 11.7W
被引数: 159
引用论文

引用论文

Binary neural networks: A survey
err2020-09-01
err303
errOAAI
errQin, Haotong; Gong, Ruihao; Liu, Xianglong; Bai, Xiao; Song, Jingkuan; Sebe, Nicu
err分享
err收藏
err分享
err收藏