arrow
Return

VLCQ: Post-training quantization for deep neural networks using variable length coding

delete2025-05-01
delete0
PRE
AI
R
Reem Abdel‐Salam
A
Ahmed H. Abdel-Gawad *
A
Amr G. Wassal
DOI:10.1016/j.future.2024.107654delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Quantization plays a crucial role inefficiently deploying deep learning models on resources constraint devices. Post-training quantization does not require either access to the original dataset or retraining the full model. Current methods that achieve high performance (near baseline results) require INT8 fixed-point integers. However, to achieve high model compression by achieving lower bit-width, significant degradation to the performance becomes the challenge. In this paper, we propose VLCQ, which relaxes the constraint of fixedpoint encoding which limits the quantization techniques from better quantizing the weights. Therefore, this work utilizes variable-length encoding which allows for exploring the whole space of quantization techniques. Thus, achieving much better results (close to or even better than the baseline results) while achieving lower bit-widths without the need to access any training data or to fine-tune the model. Extensive experiments were carried out on various deep-learning models for the image classification and segmentation, and object detection tasks. When compared to state-of-the-art post-training quantization approaches, experimental results reveal that our suggested method offers improved performance with better model compression (lower bit-rate). For per- channel quantization, our method surpassed the FP32 accuracy and Piece-Wise Linear Quantization (PWLQ) method inmost models while achieving up-to 6X model compression ratio compared to the FP32 and up-to 1.7X compared to PWLQ. If the model compression is the concern with little effect on performance, our method achieves up-to 12.25X compression ratio compared to FP32 within 4% performance loss. For per-tensor, our method is competitive with Data-Free Quantization scheme (DFQ) in achieving the best performance. However, our method is more flexible in getting lower bit rates than DFQ across the different tasks and models.
Keywords:
Post-training quantization
Deep neural networks
Non-uniform quantization
Uniform quantization
Variable length encoding
Data-free quantization
Model compression

Journal

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
Papers:
6.8K
Citations:
2.3W

Organization

E
egyptian knowledge bank (ekb)
Scholars:
11.6W
Papers: 9.3W
Citations: 84
Cited Papers

Cited Papers

errShare
errSave
Informative presence and observation in routine health data: A review of methodology for clinical risk prediction
err2020-11-09
err0
errOAAI
errRose Sisk; Lijing Lin; Matthew Sperrin; Jessica K Barrett; Brian Tom; Karla Diaz-Ordaz; Niels Peek; Glen P Martin
errShare
errSave
Loss aware post-training quantization
err2021-10-01
err54
errOAAI
errNahshan, Yury; Chmiel, Brian; Baskin, Chaim; Zheltonozhskii, Evgenii; Banner, Ron; Bronstein, Alex M.; Mendelson, Avi
errShare
errSave
Polyelectrolyte complex containing antimicrobial guanidine-based polymer and its adsorption on cellulose fibers
errhfsg
IF0
err2013-09-21
err0
PREAI
errLiying Qian; Chao Dong; Xiangtao Liang; Beihai He; Huining Xiao
errShare
errSave
errShare
errSave
researcher View more