arrow
Return

SMIT: Static mixed integer training quantization

delete2026-09-18
delete0
PRE
AI
A
Anoosh Mahjabin
H
Hakem Beitollahi *
A
Abdollah Amirkhani
P
Pejman Lotfi-Kamran
DOI:10.1007/s11227-026-08857-zdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep Neural Networks (DNNs) have achieved remarkable success in vision tasks, yet their large-scale deployment introduces substantial computational and energy demands during training. Low precision quantization offers a promising avenue for improving efficiency; however, existing sub-8-bit training techniques often encounter instability and significant accuracy degradation, limiting their practicality. This work presents Static Mixed-Integer Training (SMIT), an empirical framework for end to end DNN training using fixed, heterogeneous bit-widths for weights, activations, and gradients. SMIT is motivated by empirical observations that different training components exhibit varying sensitivity to quantization noise. Through a systematic design space exploration, we identify an empirically effective configuration W4/A6/E8 (4-bit weights, 6-bit activations, and 8-bit gradients) that provides a favorable trade off between numerical stability and reduced precision. Experimental results across multiple vision tasks, including ImageNet classification, object detection (Cityscapes), and semantic segmentation (Pascal VOC), show that SMIT consistently achieves accuracy within 1.0–2% of FP32 performance. These results demonstrate the practical feasibility of static heterogeneous sub-8-bit training while avoiding runtime precision scheduling or architecture specific tuning. Beyond numerical stability, the proposed static precision allocation may simplify future low-precision HPC implementations by avoiding runtime precision adaptation and providing a hardware-friendly framework for future low-precision accelerators and distributed AI systems.
Keywords:
Quantization
Deep neural networks (DNN)
Mixed-Integer

Journal

Journal of Supercomputing cover
Journal of Supercomputing
IF:
2.7
Papers:
1.1K
Citations:
1.0W

Organization

No organization information available
Cited Papers

Cited Papers

Distribution Adaptive INT8 Quantization for Training CNNs
err2021-05-18
err0
errOAAI
errKang Zhao; Sida Huang; Pan Pan; Yinghan Li; Yingya Zhang; Zhenyu Gu; Yinghui Xu
errShare
errSave
ImageNet Large Scale Visual Recognition Challenge
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
errShare
errSave
Training high-performance and large-scale deep neural networks with full 8-bit integers
err2020-05-01
err87
errOAAI
errYang, Yukuan; Deng, Lei; Wu, Shuang; Yan, Tianyi; Xie, Yuan; Li, Guoqi
errShare
errSave
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
err2025-01-21
err0
PREAI
errJi Lin; Jiaming Tang; Haotian Tang; Shang Yang; Guangxuan Xiao; Song Han
errShare
errSave
The Pascal Visual Object Classes (VOC) Challenge
err2009-09-09
err9.0K
PREAI
errEveringham, Mark; Van Gool, Luc; Williams, Christopher K. I.; Winn, John; Zisserman, Andrew
errShare
errSave
Is Integer Arithmetic Enough for Deep Learning Training?
err2022-01-01
err0
PREAI
errGhaffari,Alireza; Tahaei,Marzieh S.; Tayaranian,Mohammadreza; Asgharian,Masoud; Nia,Vahid Partovi
errShare
errSave
no more