返回
Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing
DOI:10.1016/j.future.2024.107698.png)
摘要
En 中文
The tensor-vector contraction (TVC) is the most memory-bound operation of its class and a core component of the higher-order power method (HOPM). This paper brings distributed-memory parallelization to a native TVC algorithm for dense tensors that overall remains oblivious to contraction mode, tensor splitting, and tensor order. Similarly, we propose a novel distributed HOPM, namely dHOPM3, that can save up to one order of magnitude of streamed memory and is about twice as costly in terms of data movement as a distributed TVC operation (dTVC) when using task-based parallelization. The numerical experiments carried out in this work on three different architectures featuring multicore and accelerators confirm that the performances of dTVC and dHOPM3 remain relatively close to the peak system memory bandwidth (50%-80%, depending on the architecture) and on par with STREAM benchmark figures. On strong scalability scenarios, our native multicore implementations of these two algorithms can achieve similar and sometimes even greater performance figures than those based upon state-of-the-art CUDA batched kernels. Finally, we demonstrate that both computation and communication can benefit from mixed precision arithmetic also incases where the hardware does not support low precision data types natively.
Keyword:
Tensor contraction
Distributed memory
High bandwidth memory
Mixed precision
GPU
Task-based parallelization
期刊
F
IF:
6.1
论文数:
6.9K
被引数:
2.3W
机构
引用论文
Heavy metal tolerance of marine phytoplankton. IV. Combined effect of zinc and cadmium on growth and uptake in some marine diatoms海洋浮游植物对重金属的耐受性。四。锌和镉对某些海洋硅藻生长和吸收的综合影响
Wide Diversity in Measurements of Growth Hormone after Stimulation Tests in Short Children are Due to Assay Variability矮小儿童刺激试验后生长激素的测量差异很大
A multi-dimensional Morton-ordered block storage for mode-oblivious tensor computations用于模式遗忘张量计算的多维Morton有序块存储

