arrow
返回

A load-balanced acceleration method for small and irregular batch matrix multiplication on GPU

delete2025-03-01
delete0
PRE
AI
Y
Yu Zhang
陆璐 (Lu Lu) *
Z
Zhanyu Yang
Z
Zhihong Liang
S
Siliang Suo
DOI:10.1016/j.sysarc.2025.103341delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
As an essential mathematical operation, GEneral Matrix Multiplication (GEMM) plays a vital role in many applications, such as high-performance computing, machine learning, etc. In practice, the performance of GEMM is limited by the dimension of matrix and the diversity of GPU hardware architectures. When dealing with batched, irregular and small matrices, the efficiency of GEMM usually performs poorly. To this end, common approach is to segment the matrix into multiple tiles and utilize parallelism between workgroups in GPU to compute the results. However, previous works only consider tile size and inter-workgroup parallelism and ignore the issues of low computational efficiency and hardware resource utilization caused by the difference in workloads between wavefronts. To address these issues, we propose a load-balanced batch GEMM acceleration method, consisting of a multi-thread kernel design and an efficient tiling algorithm. The multithread kernel design can address the workload unbalance between wavefronts indifferent workgroups, and the efficient tiling algorithm can choose the optimal tiling scheme with the new thread-level parallelism calculation method to achieve load-balanced task allocation. Finally, various comparative experiments were conducted on two GPU platforms: AMD and NVIDIA. Experimental results indicate the proposed method outperforms previous methods.
Keyword:
Batch GEMM
Thread workload
Multi-thread kernel
Tiling algorithm

期刊

Journal of Systems Architecture 封面图
Journal of Systems Architecture
IF:
4.1
论文数:
3.0K
被引数:
4.2K

机构

S
south china university of technology
学者数:
6.8W
论文数: 5.1W
被引数: 85
引用论文

引用论文

Morphology and Functional Anatomy
err2007-02-28
err0
PREAI
errJeffrey W. Shultz; Ricardo Pinto-da-Rocha
err分享
err收藏
Reconstruction of distal phalangeal injuries with the reverse homodigital island flap
err2008-12-01
err0
PREAI
errArash Momeni; Horst Zajonc; Ziad Kalash; G. Björn Stark; Holger Bannasch
err分享
err收藏
Use of explicit memory cues following parietal lobe lesions
err2012-11-01
err0
errOAAI
errIan G. Dobbins; Antonio Jaeger; Bettina Studer; Jon S. Simons
err分享
err收藏
Morphometric Comparison between Human and Rat Abducens and Oculomotor Nerves
err2008-07-16
err0
PREAI
errAttila Bardosi; Josephine Shallo; Claus Schäfer; Hermann Mühlendyck
err分享
err收藏
学者 查看更多内容