arrow
返回

Bayesian asymmetric quantized neural networks

delete2023-07-01
delete5
delete
OA
AI
J
Jen‐Tzung Chien *
DOI:10.1016/j.patcog.2023.109463delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
This paper develops a robust model compression for neural networks via parameter quantization. Tra-ditionally, quantized neural networks (QNN) were constructed by binary or ternary weights where the weights were deterministic. This paper generalizes QNN in two directions. First, M-ary QNN is developed to adjust the balance between memory storage and model capacity. The representation values and the quantization partitions in M-ary quantization are mutually estimated to enhance the resolution of gradi-ents in neural network training. A flexible quantization with asymmetric partitions is formulated. Second, the variational inference is incorporated to implement the Bayesian asymmetric QNN. The uncertainty of weights is faithfully represented to enhance the robustness of the trained model in presence of het-erogeneous data. Importantly, the multiple spike-and-slab prior is proposed to represent the quantization levels in Bayesian asymmetric learning. M-ary quantization is then optimized by maximizing the evidence lower bound of classification network. An adaptive parameter space is built to implement Bayesian quan-tization and neural representation. The experiments on various image recognition tasks show that M-ary QNN achieves similar performance as the full-precision neural network (FPNN), but the memory cost and the test time are significantly reduced relative to FPNN. The merit of Bayesian M-ary QNN using multiple spike-and-slab prior is investigated.(c) 2023 The Author(s). Published by Elsevier Ltd. This is an open access article under the CC BY license ( http://creativecommons.org/licenses/by/4.0/ )
Keyword:
Quantized neural network
Model compression
Binary neural network
Bayesian asymmetric quantization
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Pattern Recognition 封面图
Pattern Recognition
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

N
National Yang Ming Chiao Tung University
学者数:
2.5W
论文数: 2.3W
被引数: 2.2W
引用论文

引用论文

Sparse low rank factorization for deep neural network compression
err2020-07-01
err87
PREAI
errSwaminathan, Sridhar; Garg, Deepak; Kannan, Rajkumar; Andres, Frederic
err分享
err收藏
Training Multi-Bit Quantized and Binarized Networks with a Learnable Symmetric Quantizer
err2021-01-01
err11
errOAAI
errPham, Phuoc; Abraham, Jacob A.; Chung, Jaeyong
err分享
err收藏
Pruning by explaining: A novel criterion for deep neural network pruning解释修剪: 一种新的深度神经网络修剪准则
err2021-07-01
err139
errOAAI
errYeom, Seul-Ki; Seegerer, Philipp; Lapuschkin, Sebastian; Binder, Alexander; Wiedemann, Simon; Mueller, Klaus-Robert; Samek, Wojciech
err分享
err收藏
err分享
err收藏
没有更多内容