arrow
Return

Accelerating LLM Inference via Low-Bit Fine-Grained Quantization Algorithm and Bit-Level Accelerator Co-Design

delete2025-11-05
delete0
PRE
AI
X
Xilong Xie
L
L Wang
L
Limin Xiao
L
Li Ruan
T
Tairan Zhang
J
Jinquan Wang
Y
Yongyue Wang
X
Xiaojian Liao
DOI:10.1109/TC.2025.3628629delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large language models (LLMs) have emerged as one of the most impactful and transformative paradigms in natural language processing. Despite their remarkable success, the intensive computational demands and substantial memory footprint impose a significant barrier to efficient LLM inference. In this paper, we present a comprehensive solution to improve LLM inference performance under ultra-low weight precision, meticulously optimized through algorithm and architecture co-design. To achieve this, we first propose a fine-grained intra-cluster bit allocation method that partitions the weights into small clusters and explicitly considers the distribution of outliers and salient points within each cluster. Then, an intra-cluster protection mechanism is proposed to selectively preserve important weights during quantization, where an extended integer format and group-wise scale factor search are further introduced to mitigate accuracy degradation caused by aggressive bit-width reduction. Furthermore, we develop a memory-aligned encoding scheme to facilitate efficient memory access while enabling flexible identification of mixed-precision representations. Finally, we design a lightweight bit-level accelerator for low-bit LLM inference, offering simplified hardware design and enhanced adaptability through parallel bit-level computation. Compared to existing state-of-the-art quantization algorithms, our algorithm achieves higher model accuracy under ultra-low weight precision. Meanwhile, the proposed bit-level accelerator delivers speedups of 1.59$\boldsymbol{\times}$, 1.38$\boldsymbol{\times}$, and 1.61$\boldsymbol{\times}$, along with energy efficiency improvements of 1.52$\boldsymbol{\times}$, 1.42$\boldsymbol{\times}$, and 1.22$\boldsymbol{\times}$ over ANT, OliVe, and FineQ, respectively.
Keywords:
Large language models
algorithm–architecture co-design
quantization

Journal

IEEE Transactions on Computers cover
IEEE Transactions on Computers
IF:
3.8
Papers:
5.3K
Citations:
9.8K

Organization

B
beihang university
Scholars:
5.2K
Papers: 2.0K
Citations: 21
J
Jiangsu Shuguang Optoelectronics Company Ltd.
Scholars:
1
Papers: 1
Citations: 0