Return
Autonomous Model Quantization Framework for Hybrid Vision Transformers Based on Reinforcement Learning
DOI:10.1109/tcad.2025.3641538.png)
Abstract
En 中文
Existing quantization approaches often suffer from significant accuracy degradation when compressing hybrid convolution and transformer models with low bit-width. This article presents RL-PTQv2, an extension of our previous RL-PTQ framework, which introduces a new reinforcement learning (RL)-based posttraining quantization (PTQ) method. RL-PTQv2 introduces two key advances: 1) hardware (HW)-aware PTQ (optional), where RL is guided by real latency and energy feedback from an in-loop PIM simulator, enabling deployable designs that jointly optimize accuracy, latency, and energy and 2) improved quantization techniques, supporting symmetric/asymmetric quantization and mixed adaptive rounding to better balance precision and efficiency. Across various hybrid vision transformer families, including MobileViTv1 and v2, EfficientFormerv1 and v2, and MobileFormer, RL-PTQv2 achieves state-of-the-art quantized accuracy compared to previous PTQ methods. Furthermore, our quantized model showed an improvement in energy efficiency of <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$10.1\times $ </tex-math></inline-formula> on TransPIM and <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$22.6\times $ </tex-math></inline-formula> on the Titan RTX GPU compared to the baseline model, specifically when deployed on HViT-PIM, a dedicated processing framework for efficiently executing MobileViT models. HViT-PIM was developed primarily to explore the potential of HW-aware PTQ. However, the RL-PTQv2 is not limited to processing-in-memory (PIM). It can also be seamlessly integrated with a variety of bit-serial accelerators, enabling automatic quantization tailored to the underlying HW.
Keywords:
Transformer accelerator
transformer optimization
vision transformer
Journal
I
IF:
2.9
Papers:
564
Citations:
9.6K

