arrow
Return

High Performance OpenCL-Based GEMM Kernel Auto-Tuned by Bayesian Optimization

delete2025-09-01
delete0
PRE
AI
S
Shengle Lin
G
Guoqing Xiao
H
Haotian Wang
W
Wangdong Yang
李肯立 cover
李肯立 (Kenli Li)
李克勤 cover
李克勤 (Keqin Li)
DOI:10.1109/TPDS.2025.3587673delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
OpenCL has become the favored framework for emerging heterogeneous devices and FPGAs, owing to its versatility and portability. However, OpenCL-based math libraries still face challenges in fully leveraging device performance. When deploying high-performance arithmetic applications on these devices, the most important hot function is General Matrix-matrix Multiplication (GEMM). This study presents a meticulously optimized OpenCL GEMM kernel. Our enhanced GEMM kernel emphasizes two key improvements: 1) a three-level double buffer pipeline that efficiently overlaps data fetching with floating-point computations; 2) a fine-grained prefetching strategy of private memory to increase device occupancy by optimizing register unit utilization. Furthermore, this work presents a Bayesian Optimization (BO) tuner for kernel auto-tuning. Experimental results demonstrate considerable optimization improvement and performance advantages achieved on diverse OpenCL devices. Additionally, the BO tuner demonstrates superior efficiency and robustness, outperforming contemporary tuning methods.
Keywords:
OpenCL
GEMM
auto-tuning
Bayesian optimization
high-performance computing

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

H
hunan university
Scholars:
4.4W
Papers: 3.3W
Citations: 70