返回
Residual Quantization for Low Bit-Width Neural Networks
DOI:10.1109/TMM.2021.3124095.png)
摘要
En 中文
Neural network quantization has shown to be an effective way for network compression and acceleration. However, existing binary or ternary quantization methods suffer from two major issues. First, low bit-width input/activation quantization easily results in severe prediction accuracy degradation. Second, network training and quantization are always treated as two non-related tasks, leading to accumulated parameter training error and quantization error. In this work, we introduce a novel scheme, named Residual Quantization, to train a neural network with both weights and inputs constrained to low bit-width, e.g., binary or ternary values. On one hand, by recursively performing residual quantization, the resulting binary/ternary network is guaranteed to approximate the full-precision network with much smaller errors. On the other hand, we mathematically re-formulate the network training scheme in an EM-like manner, which iteratively performs network quantization and parameter optimization. During expectation, the low bit-width network is encouraged to approximate the full-precision network. During maximization, the low bit-width network is further tuned to gain better representation capability. Extensive experiments well demonstrate that the proposed quantization scheme outperforms previous low bit-width methods and achieves much closer performance to the full-precision counterpart.
Keyword:
Quantization (signal)
Training
Computational modeling
Neurons
Degradation
Task analysis
Optimization
Deep learning
network quantization
binarization
network acceleration
期刊
IF:
9.7
论文数:
4.5K
被引数:
2.4W
机构
引用论文
A Hierarchical Cardiac Rhythm Classification Methodology Based on Electrocardiogram Fiducial Points一种基于心电图基准点的分层心律分类方法
A Multi-Task Neural Approach for Emotion Attribution, Classification, and Summarization用于情感归因,分类和总结的多任务神经方法

