返回
General Bitwidth Assignment for Efficient Deep Convolutional Neural Network Quantization
DOI:10.1109/TNNLS.2021.3069886.png)
摘要
En 中文
Model quantization is essential to deploy deep convolutional neural networks (DCNNs) on resource-constrained devices. In this article, we propose a general bitwidth assignment algorithm based on theoretical analysis for efficient layerwise weight and activation quantization of DCNNs. The proposed algorithm develops a prediction model to explicitly estimate the loss of classification accuracy led by weight quantization with a geometrical approach. Consequently, dynamic programming is adopted to achieve optimal bitwidth assignment on weights based on the estimated error. Furthermore, we optimize bitwidth assignment for activations by considering the signal-to-quantization-noise ratio (SQNR) between weight and activation quantization. The proposed algorithm is general to reveal the tradeoff between classification accuracy and model size for various network architectures. Extensive experiments demonstrate the efficacy of the proposed bitwidth assignment algorithm and the error rate prediction model. Furthermore, the proposed algorithm is shown to be well extended to object detection.
Keyword:
Quantization (signal)
Prediction algorithms
Error analysis
Computational modeling
Training
Convolutional neural networks
Standards
Bitwidth assignment
deep convolutional neural networks (DCNNs)
network quantization
quantization noise
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
8.9
论文数:
7.6K
被引数:
7.2W
机构
引用论文
A Hierarchical Cardiac Rhythm Classification Methodology Based on Electrocardiogram Fiducial Points一种基于心电图基准点的分层心律分类方法
Neural Network Training With Levenberg-Marquardt and Adaptable Weight Compression基于levenberg-marquardt和自适应权值压缩的神经网络训练

