arrow
Return

Autonomous Model Quantization Framework for Hybrid Vision Transformers Based on Reinforcement Learning

delete2025-12-08
delete0
PRE
AI
E
Eunji Kwon
T
Tajana Rosing
DOI:10.1109/tcad.2025.3641538delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Existing quantization approaches often suffer from significant accuracy degradation when compressing hybrid convolution and transformer models with low bit-width. This article presents RL-PTQv2, an extension of our previous RL-PTQ framework, which introduces a new reinforcement learning (RL)-based posttraining quantization (PTQ) method. RL-PTQv2 introduces two key advances: 1) hardware (HW)-aware PTQ (optional), where RL is guided by real latency and energy feedback from an in-loop PIM simulator, enabling deployable designs that jointly optimize accuracy, latency, and energy and 2) improved quantization techniques, supporting symmetric/asymmetric quantization and mixed adaptive rounding to better balance precision and efficiency. Across various hybrid vision transformer families, including MobileViTv1 and v2, EfficientFormerv1 and v2, and MobileFormer, RL-PTQv2 achieves state-of-the-art quantized accuracy compared to previous PTQ methods. Furthermore, our quantized model showed an improvement in energy efficiency of <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$10.1\times $ </tex-math></inline-formula> on TransPIM and <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$22.6\times $ </tex-math></inline-formula> on the Titan RTX GPU compared to the baseline model, specifically when deployed on HViT-PIM, a dedicated processing framework for efficiently executing MobileViT models. HViT-PIM was developed primarily to explore the potential of HW-aware PTQ. However, the RL-PTQv2 is not limited to processing-in-memory (PIM). It can also be seamlessly integrated with a variety of bit-serial accelerators, enabling automatic quantization tailored to the underlying HW.
Keywords:
Transformer accelerator
transformer optimization
vision transformer

Journal

I
IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems
IF:
2.9
Papers:
564
Citations:
9.6K

Organization

U
university of california san diego
Scholars:
5.1K
Papers: 2.3K
Citations: 1
K
Kookmin University
Scholars:
457
Papers: 246
Citations: 3.3K