返回
Bayesian asymmetric quantized neural networks
DOI:10.1016/j.patcog.2023.109463.png)
摘要
En 中文
This paper develops a robust model compression for neural networks via parameter quantization. Tra-ditionally, quantized neural networks (QNN) were constructed by binary or ternary weights where the weights were deterministic. This paper generalizes QNN in two directions. First, M-ary QNN is developed to adjust the balance between memory storage and model capacity. The representation values and the quantization partitions in M-ary quantization are mutually estimated to enhance the resolution of gradi-ents in neural network training. A flexible quantization with asymmetric partitions is formulated. Second, the variational inference is incorporated to implement the Bayesian asymmetric QNN. The uncertainty of weights is faithfully represented to enhance the robustness of the trained model in presence of het-erogeneous data. Importantly, the multiple spike-and-slab prior is proposed to represent the quantization levels in Bayesian asymmetric learning. M-ary quantization is then optimized by maximizing the evidence lower bound of classification network. An adaptive parameter space is built to implement Bayesian quan-tization and neural representation. The experiments on various image recognition tasks show that M-ary QNN achieves similar performance as the full-precision neural network (FPNN), but the memory cost and the test time are significantly reduced relative to FPNN. The merit of Bayesian M-ary QNN using multiple spike-and-slab prior is investigated.(c) 2023 The Author(s). Published by Elsevier Ltd. This is an open access article under the CC BY license ( http://creativecommons.org/licenses/by/4.0/ )
Keyword:
Quantized neural network
Model compression
Binary neural network
Bayesian asymmetric quantization
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.6
论文数:
1.3W
被引数:
4.5W
机构
引用论文
Training Multi-Bit Quantized and Binarized Networks with a Learnable Symmetric Quantizer
IEEE ACCESS
IF3.6
Training high-performance and large-scale deep neural networks with full 8-bit integers使用完整的8位整数训练高性能和大规模的深度神经网络
NEURAL NETWORKS
IF6.3
Pruning by explaining: A novel criterion for deep neural network pruning解释修剪: 一种新的深度神经网络修剪准则
PATTERN RECOGNITION
IF7.6
没有更多内容

