arrow
返回

Learning Generalizable Mixed-Precision Quantization via Attribution Imitation

delete2024-06-02
delete0
PRE
AI
Z
Ziwei Wang
韩
韩笑 (Xiao Han)
J
Jie Zhou
J
Jiwen Lu *
DOI:10.1007/s11263-024-02130-7delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In this paper, we propose a generalizable mixed-precision quantization (GMPQ) method for efficient inference. Conventional methods require the consistency of datasets for bitwidth search and model deployment to guarantee the policy optimality, leading to heavy search cost on challenging large-scale datasets in realistic applications. On the contrary, our GMPQ searches the mixed-quantization policy that can be generalized to large-scale datasets with only a small amount of data, so that the search cost is significantly reduced without performance degradation. Specifically, we observe that locating network attribution correctly is general ability for accurate visual analysis across different data distribution. Therefore, despite of pursuing higher accuracy and lower model complexity, we preserve attribution rank consistency between the quantized models and their full-precision counterparts via capacity-aware attribution imitation for generalizable mixed-precision quantization strategy search, where the capacity of quantized networks is considered to fully utilize the network capacity without insufficiency. Since slight noise in attribution is amplified by discrete ranking operations with significant rank errors, mimicking the attribution ranks of the full-precision models obstructs the quantized networks to correctly locate the attribution. To address this, we further present a robust generalizable mixed-precision quantization method to smooth the attribution for rank error alleviation by hierarchical attribution partitioning, which efficiently partitions the attribution pixels in high spatial resolution and assigns the same attribution value for pixels within a group. Moreover, we propose dynamic capacity-aware attribution imitation to adjust the concentration degree of the attribution according to sample hardness, so that sufficient model capacity is achieved with full utilization for each image. Extensive experiments on image classification and object detection show that our GMPQ and R-GMPQ obtain competitive accuracy-complexity trade-offs with significantly reduced search cost compared to the state-of-the-art mixed-precision networks.
Keyword:
Mixed-precision quantization
Generalizable compression policy
Attribution rank preservation
Attribution imitation
Hierarchical attribution partitioning

期刊

International Journal of Computer Vision 封面图
International Journal of Computer Vision
IF:
9.3
论文数:
3.9K
被引数:
2.8W

机构

T
tsinghua university
学者数:
11.9W
论文数: 10.0W
被引数: 137
引用论文

引用论文

An analysis of Recurrent Neural Networks for Botnet detection behavior
err2016-06-01
err0
PREAI
errPablo Torres; Carlos Catania; Sebastian Garcia; Carlos Garcia Garino
err分享
err收藏
Client ahead‐of‐time compiler for embedded Java platforms
err2008-08-05
err0
PREAI
errSunghyun Hong; Jin‐Chul Kim; Soo‐Mook Moon; Jin Woo Shin; Jaemok Lee; Hyeong‐Seok Oh; Hyung‐Kyu Choi
err分享
err收藏
High prevalence of malaria in a non-endemic setting among febrile episodes in travellers and migrants coming from endemic areas: a retrospective analysis of a 2013–2018 cohort
err2021-11-27
err0
errOAAI
errAlejandro Garcia-Ruiz de Morales; Covadonga Morcate; Elena Isaba-Ares; Ramon Perez-Tanoira; Jose A. Perez-Molina
err分享
err收藏
err分享
err收藏
学者 查看更多内容