Return
QuantTPM: Efficient Mixed-Precision Quantization Framework for Tractable Probabilistic Models
DOI:10.1109/TCAD.2025.3543424.png)
Abstract
En 中文
Tractable probabilistic models (TPMs) can perform reliable probabilistic inference and enhance the reasoning capabilities of edge devices, such as aiding decision-making for autonomous vehicles. To deploy TPMs in edge scenarios with constrained hardware resources and energy, efficient quantization algorithms are necessary. However, the traditional quantization methods for neural networks are not applicable to TPMs due to the irregular model structure and highly varying data distribution. To address the issues, we propose QuantTPM, a mixed-precision quantization framework designed to enhance the energy and resource efficiency of TPM inference. First, we reformulate the irregular model structure into a unified format, as irregular structures are inefficient for hardware implementation. Second, we divide the reformulated model graph into hierarchical levels, so as to assign appropriate quantization bit-widths for different levels with varying precision requirements. Third, we decompose the entire mixed-precision quantization search into several steps with smaller search spaces, so as to reduce the algorithm complexity and save search time. Compared with state-of-the-art works, our mixed-precision quantization framework achieves, on average, $3.7\times $ weight compression, $6.0\times $ resource efficiency, and $4.8\times $ energy consumption, while maintaining competitive accuracy.
Keywords:
Mixed precision quantization
probabilistic inference
sum-product networks
tractable probabilistic model (TPM)
Journal
I
IF:
2.9
Papers:
586
Citations:
9.6K

