arrow
Return

An efficient quantized GEMV implementation for large language models inference with matrix core

delete2025-02-14
delete0
PRE
AI
张宇 (Yu Zhang)
陆璐 (Lu Lu) *
R
Rong Zhao
Y
Yijie Guo
Z
Zhanyu Yang
DOI:10.1007/s11227-025-06993-6delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The impressive advantages of Large Language Models (LLMs) have sparked much attention in deploying and utilizing these models on devices. However, the excessive parameters in LLMs lead to a significant memory footprint and computing burden during the inference process, severely restricting the potential uses of LLMs. As an effective model compression method, quantized compression can lower the threshold for deployment and inference of LLMs. In this way, quantized GEneral Matrix-Vector multiplication (GEMV) is the primary runtime component in the inference process. In practice, the dequantization process and low computational density limit the performance of quantized GEMV. This paper proposes an efficient quantized GEMV implementation consisting of the vectorized pre-fetch scheme, an efficient kernel design based on Matrix Core, and the optimization of atomicAdd to accelerate the LLMs' inference. Finally, various comparison experiments were performed on MI210. Experiment results show the proposed method performs better than previous approaches in multiple shape quantized GEMV and end-to-end inference on LLMs.
Keywords:
GPU
LLMs
Quantized GEMV
Matrix core

Journal

Journal of Supercomputing cover
Journal of Supercomputing
IF:
2.7
Papers:
1.1K
Citations:
1.0W

Organization

No organization information available