返回
Exploring the Design Space of Distributed Parallel Sparse Matrix-Multiple Vector Multiplication
DOI:10.1109/TPDS.2024.3452478.png)
摘要
En 中文
We consider the distributed memory parallel multiplication of a sparse matrix by a dense matrix (SpMM). The dense matrix is often a collection of dense vectors. Standard implementations will multiply the sparse matrix by multiple dense vectors at the same time, to exploit the computational efficiencies therein. But such approaches generally utilize the same sparse matrix partitioning as if multiplying by a single vector. This article explores the design space of parallelizing SpMM and shows that a coarser grain partitioning of the matrix combined with a column-wise partitioning of the block of vectors can often require less communication volume and achieve higher SpMM performance. An algorithm is presented that chooses a process grid geometry for a given number of processes to optimize the performance of parallel SpMM. The algorithm can augment existing graph partitioners by utilizing the additional concurrency available when multiplying by multiple dense vectors to further reduce communication.
Keyword:
Sparse matrices
Partitioning algorithms
Vectors
Costs
Three-dimensional displays
Space exploration
Optimization
SpMM
SpMV
distributed-memory matrix multiplication
communication optimization
期刊
IF:
6
论文数:
5.2K
被引数:
1.1W
机构
引用论文
A two-dimensional data distribution method for parallel sparse matrix-vector multiplication
SIAM REVIEW
IF6.1
A 3-Hydroxypropionate/4-Hydroxybutyrate Autotrophic Carbon Dioxide Assimilation Pathway in Archaea
Science
IF0
MPI-FAUN: An MPI-Based Framework for Alternating-Updating Nonnegative Matrix FactorizationMpi-faun: 基于MPI的交替更新非负矩阵分解框架

