返回
A Novel Parallel Algorithm for Sparse Tensor Matrix Chain Multiplication via TCU-Acceleration
DOI:10.1109/TPDS.2023.3288520.png)
摘要
En 中文
Analysis of multi-dimensional data, especially tensor decomposition, which extracts latent information, is becoming considerably popular. Although multi-dimensional sparse data is typically processed on multi-core processors, developing highly optimized GPU-based Sparse Tensor Matrix Chain Multiplication (SpTMCM) is challenging. The purpose of this paper is to investigate a novel approach named SpTMCM and to explore the discovery of SpTMCM coupled with the emerging computing core, Tensor Core Unit (TCU). In contrast to prior work, the proposed novel approach enables a uniform storage format and optimization approach for SpTMCM. We design a hybrid tensor format based on multi-dimensional tiling that divides the tensor depending on the tile threshold to address the inefficient memory accesses caused by the irregular nonzero distribution of the sparse tensor. Further, we develop a TCU-based tensor parallel algorithm with our novel approach to increase the memory bandwidth. Compared to stateof-the-art works, our method achieves 1.16 similar to 24.12x speedup for SpMTTKRP and 5.07 similar to 7.15x speedup for SpTTMChain across NVIDIA A100 GPU on a range of real-world sparse tensors.
Keyword:
GPU
hybrid format
parallel performance
SpMTTKRP
SpTTMChain
sparse tensor
tensor core
期刊
IF:
6
论文数:
5.2K
被引数:
1.1W
机构
引用论文
True Load Balancing for Matricized Tensor Times Khatri-Rao ProductMatricized Tensor Times khatri-rao产品的真正负载平衡

