arrow
Return

Optimizing sparse-dense matrix-matrix multiplication for DCUs

delete2025-11-01
delete0
PRE
AI
郭恒亮 (Hengliang Guo)
Y
Yubo Han
H
Haolei Wang
S
Shengguang Zhu
G
Gang Wu
Y
Yang Guo *
刘向东 cover
刘向东 (Xiangdong Liu)
C
C.D. Li
DOI:10.1007/s42514-025-00254-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
To address the issues of sparse matrix load imbalance and parallelism degradation with increasing matrix size in the mainstream Sparse-dense matrix-matrix multiplication (SpMM) parallelization strategy row-split, we propose a new framework for parallel SpMM computation on DCUs (GPU-like accelerators). This framework is based on the standard CSR format, requiring no additional format conversion, and thus offers strong generality. To address the issue of load imbalance, we introduce a coarse-grained two-level binning strategy that categorizes the rows of the sparse matrix into three groups based on the number of non-zero elements. Dedicated computation kernels are designed for each category to better accommodate different types of computational tasks, thereby significantly improving load balance. To address the decline in parallelism as the matrix size increases, we design multiple optimized kernels and dynamically select the optimal configuration at runtime to maximize parallelism. Experimental results show that our proposed SpMM framework significantly outperforms two current state-of-the-art row-split based SpMM algorithms (rocSparse and GE-SpMM), achieving speedups of 5.4\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\times $$\end{document} and 2.28\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\times $$\end{document}, respectively.
Keywords:
SpMM
Sparse matrix
Load balancing
DCU accelerator

Journal

C
CCF Transactions on High Performance Computing
IF:
1.9
Papers:
38
Citations:
253

Organization

S
sinopec
Scholars:
602
Papers: 283
Citations: 0
Z
Zhengzhou University
Scholars:
6.8W
Papers: 4.4W
Citations: 8.5W