arrow
返回

Deployable mixed-precision quantization with co-learning and one-time search

delete2025-01-01
delete0
PRE
AI
S
Shiguang Wang
Z
Zhongyu Zhang
G
Guo Ai
J
Jian Cheng *
DOI:10.1016/j.neunet.2024.106812delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Mixed-precision quantization plays a pivotal role in deploying deep neural networks in resource-constrained environments. However, the task of finding the optimal bit-width configurations for different layers under deployable mixed-precision quantization has barely been explored and remains a challenge. In this work, we present Cobits, an efficient and effective deployable mixed-precision quantization framework based on the relationship between the range of real-valued input and the range of quantized real-valued. It assigns a higher bit-width to the quantizer with a narrower quantized real-valued range and a lower bit-width to the quantizer with a wider quantized real-valued range. Cobits employs a co-learning approach to entangle and learn quantization parameters across various bit-widths, distinguishing between shared and specific parts. The shared part collaborates, while the specific part isolates precision conflicts. Additionally, we upgrade the normal quantizer to dynamic quantizer to mitigate statistical issues in the deployable mixed-precision supernet. Over the trained mixed-precision supernet, we utilize the quantized real-valued ranges to derive quantized- bit-sensitivity, which can serve as importance indicators for efficiently determining bit-width configurations, eliminating the need for iterative validation dataset evaluations. Extensive experiments show that Cobits outperforms previous state-of-the-art quantization methods on the ImageNet and COCO datasets while retaining superior efficiency. We show this approach dynamically adapts to varying bit-width and can generalize to various deployable backends. The code will be made public in https://github.com/sunnyxiaohu/cobits.
Keyword:
Model quantization
Mixed-precision quantization
Deployable quantization
Hardware quantization
Model compression

期刊

Neural Networks 封面图
Neural Networks
IF:
6.3
论文数:
7.8K
被引数:
3.0W

机构

暂无机构信息
引用论文

引用论文

err分享
err收藏
Thermal Plasmas
err
IF0
err1994-01-01
err0
PREAI
errMaher I. Boulos; Pierre Fauchais; Emil Pfender
err分享
err收藏
err分享
err收藏
Depression Trajectories of Antenatally Depressed and Nondepressed Young Mothers: Implications for Child Socioemotional Development
err2016-05-01
err0
PREAI
errMaryna Raskin; M. Ann Easterbrooks; Renee S. Lamoreau; Chie Kotake; Jessica Goldberg
err分享
err收藏
学者 查看更多内容