arrow
返回

Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing

delete2025-05-01
delete0
PRE
AI
P
Pedro J. Martínez-Ferrer *
A
Albert-Jan N. Yzelman
V
Vicenç Beltrán
DOI:10.1016/j.future.2024.107698delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The tensor-vector contraction (TVC) is the most memory-bound operation of its class and a core component of the higher-order power method (HOPM). This paper brings distributed-memory parallelization to a native TVC algorithm for dense tensors that overall remains oblivious to contraction mode, tensor splitting, and tensor order. Similarly, we propose a novel distributed HOPM, namely dHOPM3, that can save up to one order of magnitude of streamed memory and is about twice as costly in terms of data movement as a distributed TVC operation (dTVC) when using task-based parallelization. The numerical experiments carried out in this work on three different architectures featuring multicore and accelerators confirm that the performances of dTVC and dHOPM3 remain relatively close to the peak system memory bandwidth (50%-80%, depending on the architecture) and on par with STREAM benchmark figures. On strong scalability scenarios, our native multicore implementations of these two algorithms can achieve similar and sometimes even greater performance figures than those based upon state-of-the-art CUDA batched kernels. Finally, we demonstrate that both computation and communication can benefit from mixed precision arithmetic also incases where the hardware does not support low precision data types natively.
Keyword:
Tensor contraction
Distributed memory
High bandwidth memory
Mixed precision
GPU
Task-based parallelization

期刊

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
论文数:
6.9K
被引数:
2.3W

机构

B
barcelona supercomputer center (bsc-cns)
学者数:
1.2K
论文数: 825
被引数: 5
U
universitat politecnica de catalunya
学者数:
1.9W
论文数: 1.6W
被引数: 17
引用论文

引用论文

err分享
err收藏
Tensor Decompositions and Applications张量分解及其应用
err2009-08-05
err7.5K
PREAI
errKolda, Tamara G.; Bader, Brett W.
err分享
err收藏
A massively parallel tensor contraction framework for coupled-cluster computations
err2014-12-01
err148
PREAI
errSolomonik, Edgar; Matthews, Devin; Hammond, Jeff R.; Stanton, John F.; Demmel, James
err分享
err收藏
Intermolecular 2+2 Imine-Olefin Photocycloadditions Enabled by Cu(I)-Alkene MLCT
err
IF0
err2021-06-28
err0
errOAAI
errDaniel Flores; Michael Neville; Valerie Schmidt
err分享
err收藏
学者 查看更多内容