arrow
Return

cuFastTucker-2L: A Two-Level Optimization Parallel Algorithm for Solving FastTucker Decomposition on GPU Platform

delete2026-06-22
delete0
PRE
AI
Z
Zixuan Li
H
H P Wang
W
Wangdong Yang
李克勤 cover
李克勤 (Keqin Li)
李肯立 cover
李肯立 (Kenli Li)
DOI:10.1109/tpds.2026.3706077delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
High-order, High-dimensional, Large-scale Sparse Tensors (HHLST) are widely utilized across various scientific research fields, making direct analysis and learning of HHLST increasingly impractical. Tensor decomposition techniques provide a powerful approach for analyzing and learning HHLST; however, existing parallel tensor decomposition libraries are unable to effectively handle HHLST. While the introduction of FastTucker decomposition alleviates the significant computational and memory overhead associated with Tucker decomposition, the parallelism of current FastTucker decomposition algorithms remains constrained by the rank of the decomposition. To address these limitations, this paper proposes a novel two-level parallel FastTucker decomposition algorithm, FastTucker-2 L, which supports larger-rank FastTucker decomposition for HHLST. Furthermore, an efficient parallel sparse FastTucker-2 L algorithm, cuFastTucker-2 L, is implemented on the CUDA GPU platform. The upper-level optimization in cuFastTucker-2 L focuses on sparse CP decomposition. Experimental results demonstrate that the upper-level optimization in the FastTucker-2 L algorithm achieves a <inline-formula><tex-math notation="LaTeX">$5.00\times$</tex-math></inline-formula> to <inline-formula><tex-math notation="LaTeX">$175.00\times$</tex-math></inline-formula> speedup compared to state-of-the-art sparse CP decomposition algorithms, while the overall FastTucker-2L algorithm delivers a <inline-formula><tex-math notation="LaTeX">$1.98\times$</tex-math></inline-formula> to <inline-formula><tex-math notation="LaTeX">$2.45\times$</tex-math></inline-formula> speedup compared to the state-of-the-art sparse FastTucker decomposition algorithms.
Keywords:
GPU CUDA parallelization
CP decomposition
tucker decomposition
FastTucker decomposition
sparse tensor decomposition
tensor computation

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

H
hunan university
Scholars:
4.4W
Papers: 3.3W
Citations: 70
X
xiangtan university
Scholars:
1.5W
Papers: 9.1K
Citations: 8