Return
SMIT: Static mixed integer training quantization
DOI:10.1007/s11227-026-08857-z.png)
Abstract
En 中文
Deep Neural Networks (DNNs) have achieved remarkable success in vision tasks, yet their large-scale deployment introduces substantial computational and energy demands during training. Low precision quantization offers a promising avenue for improving efficiency; however, existing sub-8-bit training techniques often encounter instability and significant accuracy degradation, limiting their practicality. This work presents Static Mixed-Integer Training (SMIT), an empirical framework for end to end DNN training using fixed, heterogeneous bit-widths for weights, activations, and gradients. SMIT is motivated by empirical observations that different training components exhibit varying sensitivity to quantization noise. Through a systematic design space exploration, we identify an empirically effective configuration W4/A6/E8 (4-bit weights, 6-bit activations, and 8-bit gradients) that provides a favorable trade off between numerical stability and reduced precision. Experimental results across multiple vision tasks, including ImageNet classification, object detection (Cityscapes), and semantic segmentation (Pascal VOC), show that SMIT consistently achieves accuracy within 1.0–2% of FP32 performance. These results demonstrate the practical feasibility of static heterogeneous sub-8-bit training while avoiding runtime precision scheduling or architecture specific tuning. Beyond numerical stability, the proposed static precision allocation may simplify future low-precision HPC implementations by avoiding runtime precision adaptation and providing a hardware-friendly framework for future low-precision accelerators and distributed AI systems.
Keywords:
Quantization
Deep neural networks (DNN)
Mixed-Integer
Journal
IF:
2.7
Papers:
1.1K
Citations:
1.0W
Organization
No organization information available
Cited Papers
Training high-performance and large-scale deep neural networks with full 8-bit integers
NEURAL NETWORKS
IF6.3
no more

