Return
I-Segmenter: Integer-only Vision Transformer for efficient semantic segmentation
J
M
M
DOI:10.1016/j.sysarc.2026.103888.png)
Abstract
En 中文
• We propose I-Segmenter, which, to the best of our knowledge, is the first ViT-based semantic segmentation model supporting full integer-only inference. • We introduce λ -ShiftGELU, a novel integer-compatible activation function that stabilizes QAT and PTQ for dense prediction tasks. • We remove the L2 normalization step and replace the floating-point bilinear upsampling with nearest neighbor interpolation, a lightweight and integer-friendly alternative widely supported in inference engines, and perform an ablation study to assess their impact on segmentation accuracy. • Through a comprehensive evaluation across datasets, model scales, and input sizes, we show that I-Segmenter maintains accuracy within a negligible margin of its FP32 counterpart (no more than 6.7% loss), while reducing model size by up to 3.8× and achieving up to 1.2× faster inference with optimized runtimes. • We evaluate I-Segmenter in the one-shot PTQ setting, where quantization is performed using a single calibration image. This approach is highly practical but also challenging. Remarkably, we achieve competitive accuracy with just 1 s of calibration time. • We analyze the role of inference backends (ONNX Runtime, TensorRT, and TVM) in exploiting integer-only execution, highlighting both performance benefits and engineering challenges for deployment on edge devices. By leveraging TVM, we demonstrate that I-Segmenter can be executed entirely with integer arithmetic, down to the level of individual kernel computations.
Journal
IF:
4.1
Papers:
2.9K
Citations:
4.2K
Organization
No organization information available
