arrow
Return

Accelerating Sparse Transformer Inference on GPU

delete2026-01-01
delete0
PRE
AI
W
Wenhao Dai
H
Haodong Deng
M
Mengfei Rong
X
Xinyu Yang
H
Hongyu Liu
F
Fangxin Liu
Y
Yang, Hailong
Q
Qianwen Cao
Q
Qingxiao Sun *
DOI:10.1145/3774934.3786434delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large language models (LLMs) are popular around the world due to their powerful understanding capabilities. As the core component of LLMs, accelerating Transformer through parallelization has gradually become a hot research topic. Mask layers introduce sparsity into Transformer to reduce calculations. However, previous works rarely focus on the performance optimization of sparse Transformer. In addition, current static operator fusion schemes fail to adapt to diverse application scenarios. To address the above problems, we propose STOF, a framework that incorporates optimizations for Sparse Transformer that enables flexible masking and Operator Fusion on GPU. For multi-head attention (MHA) structure, STOF maps the computation to row-wise or blockwise kernels with unique storage formats according to analytical modeling. For downstream operators, STOF maps the fusion scheme to compilation templates and determines the optimal running configuration through two-stage searching.The experimental results show that compared to the stateof-the-art work, STOF achieves maximum speedups of 1.6x in MHA computation and 1.4x in end-to-end inference.
Keywords:
GPU
Sparse Transformer
Multi-head Attention
Operator Fusion

Journal

P
PROCEEDINGS OF THE 31ST ACM SIGPLAN ANNUAL SYMPOSIUM ON PRINCIPLES AND PRACTICE OF PARALLEL PROGRAMMING, PPOPP 2026
IF:
0
Papers:
43
Citations:
0

Organization

B
Beihang University
Scholars:
5.1W
Papers: 4.1W
Citations: 37
S
shanghai jiao tong university
Scholars:
15.5W
Papers: 11.6W
Citations: 159
B
baidu
Scholars:
577
Papers: 470
Citations: 1
C
china university of petroleum
Scholars:
4.1W
Papers: 2.7W
Citations: 30
researcher View more organizations