arrow
Return

QuantTPM: Efficient Mixed-Precision Quantization Framework for Tractable Probabilistic Models

delete2025-01-01
delete0
PRE
AI
张申 (Shen Zhang)
B
Bin Ning
G
Guangyao Yan
X
Xinzhe Liu
W
Weixiong Jiang
Y
Yajun Ha
DOI:10.1109/TCAD.2025.3543424delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Tractable probabilistic models (TPMs) can perform reliable probabilistic inference and enhance the reasoning capabilities of edge devices, such as aiding decision-making for autonomous vehicles. To deploy TPMs in edge scenarios with constrained hardware resources and energy, efficient quantization algorithms are necessary. However, the traditional quantization methods for neural networks are not applicable to TPMs due to the irregular model structure and highly varying data distribution. To address the issues, we propose QuantTPM, a mixed-precision quantization framework designed to enhance the energy and resource efficiency of TPM inference. First, we reformulate the irregular model structure into a unified format, as irregular structures are inefficient for hardware implementation. Second, we divide the reformulated model graph into hierarchical levels, so as to assign appropriate quantization bit-widths for different levels with varying precision requirements. Third, we decompose the entire mixed-precision quantization search into several steps with smaller search spaces, so as to reduce the algorithm complexity and save search time. Compared with state-of-the-art works, our mixed-precision quantization framework achieves, on average, $3.7\times $ weight compression, $6.0\times $ resource efficiency, and $4.8\times $ energy consumption, while maintaining competitive accuracy.
Keywords:
Mixed precision quantization
probabilistic inference
sum-product networks
tractable probabilistic model (TPM)

Journal

I
IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems
IF:
2.9
Papers:
586
Citations:
9.6K

Organization

S
ShanghaiTech University
Scholars:
9.6K
Papers: 5.9K
Citations: 1.6W