arrow
返回

VLCQ: Post-training quantization for deep neural networks using variable length coding

delete2025-05-01
delete0
PRE
AI
R
Reem Abdel‐Salam
A
Ahmed H. Abdel-Gawad *
A
Amr G. Wassal
DOI:10.1016/j.future.2024.107654delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Quantization plays a crucial role inefficiently deploying deep learning models on resources constraint devices. Post-training quantization does not require either access to the original dataset or retraining the full model. Current methods that achieve high performance (near baseline results) require INT8 fixed-point integers. However, to achieve high model compression by achieving lower bit-width, significant degradation to the performance becomes the challenge. In this paper, we propose VLCQ, which relaxes the constraint of fixedpoint encoding which limits the quantization techniques from better quantizing the weights. Therefore, this work utilizes variable-length encoding which allows for exploring the whole space of quantization techniques. Thus, achieving much better results (close to or even better than the baseline results) while achieving lower bit-widths without the need to access any training data or to fine-tune the model. Extensive experiments were carried out on various deep-learning models for the image classification and segmentation, and object detection tasks. When compared to state-of-the-art post-training quantization approaches, experimental results reveal that our suggested method offers improved performance with better model compression (lower bit-rate). For per- channel quantization, our method surpassed the FP32 accuracy and Piece-Wise Linear Quantization (PWLQ) method inmost models while achieving up-to 6X model compression ratio compared to the FP32 and up-to 1.7X compared to PWLQ. If the model compression is the concern with little effect on performance, our method achieves up-to 12.25X compression ratio compared to FP32 within 4% performance loss. For per-tensor, our method is competitive with Data-Free Quantization scheme (DFQ) in achieving the best performance. However, our method is more flexible in getting lower bit rates than DFQ across the different tasks and models.
Keyword:
Post-training quantization
Deep neural networks
Non-uniform quantization
Uniform quantization
Variable length encoding
Data-free quantization
Model compression

期刊

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
论文数:
6.8K
被引数:
2.3W

机构

E
egyptian knowledge bank (ekb)
学者数:
11.6W
论文数: 9.3W
被引数: 84
引用论文

引用论文

err分享
err收藏
Loss aware post-training quantization损失感知训练后量化
err2021-10-01
err54
errOAAI
errNahshan, Yury; Chmiel, Brian; Baskin, Chaim; Zheltonozhskii, Evgenii; Banner, Ron; Bronstein, Alex M.; Mendelson, Avi
err分享
err收藏
Polyelectrolyte complex containing antimicrobial guanidine-based polymer and its adsorption on cellulose fibers
errhfsg
IF0
err2013-09-21
err0
PREAI
errLiying Qian; Chao Dong; Xiangtao Liang; Beihai He; Huining Xiao
err分享
err收藏
err分享
err收藏
学者 查看更多内容