返回
PARALLEL ALGORITHMS FOR TENSOR TRAIN ARITHMETIC
DOI:10.1137/20M1387158.png)
摘要
En 中文
We present efficient and scalable parallel algorithms for performing mathematical operations for low-rank tensors represented in the tensor train (TT) format. We consider algorithms for addition, elementwise multiplication, computing norms and inner products, orthonormalization, and rounding (rank truncation). These are the kernel operations for applications such as iterative Krylov solvers that exploit the TT structure. The parallel algorithms are designed for distributed-memory computation, and we propose a data distribution and strategy that parallelizes computations for individual cores within the TT format. We analyze the computation and communication costs of the proposed algorithms to show their scalability, and we present numerical experiments that demonstrate their efficiency on both shared-memory and distributed-memory parallel systems. For example, we observe better single-core performance than the existing MATLAB TT-Toolbox in rounding a 2GB TT tensor, and our implementation achieves a 34x speedup using all 40 cores of a single node. We also show nearly linear parallel scaling on larger TT tensors up to over 10,000 cores for all mathematical operations.
Keyword:
low-rank tensor format
tensor train
parallel algorithms
QR
SVD
期刊
IF:
2.6
论文数:
5.1K
被引数:
1.8W
机构
引用论文
A Fractional-Order Memristive Two-Neuron-Based Hopfield Neuron Network: Dynamical Analysis and Application for Image Encryption
Mathematics
IF0
Heavy metal tolerance of marine phytoplankton. IV. Combined effect of zinc and cadmium on growth and uptake in some marine diatoms海洋浮游植物对重金属的耐受性。四。锌和镉对某些海洋硅藻生长和吸收的综合影响

