Return
Accelerating MTTKRP for Sparse Tensor Decomposition on GPUs
DOI:10.1016/j.jpdc.2026.105239.png)
Abstract
En 中文
• We introduce a novel parallel algorithm to perform MTTKRP on sparse tensors using multiple GPUs. The proposed algorithm achieves geometric mean speedups of 1.7x, 2.1x, and 2.4x in execution time while using 2, 3, and 4 GPUs, compared to a single GPU implementation on realworld tensors. • We introduce a load balancing scheme to distribute the computations of MTTKRP among GPUs while eliminating the communication of intermediate values across GPUs in each mode of computation. After each mode of computation, the updated factor matrix rows are shared across the GPUs. The proposed load balancing scheme shows less than 1 computation time across GPUs. • Our work achieves a geometric mean speedup of 3.0x in total execution time compared to stateof-the-art on real-world tensors from a diverse set of applications.
Keywords:
MTTKRP
sparse tensor decomposition
GPU acceleration
parallel algorithm
load balancing
Journal
IF:
4
Papers:
3.8K
Citations:
4.8K

