arrow
返回

A Native Tensor-Vector Multiplication Algorithm for High Performance Computing

delete2022-12-01
delete1
delete
OA
AI
P
Pedro J. Martínez-Ferrer *
V
Vicenç Beltrán
DOI:10.1109/TPDS.2022.3153113delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Tensor computations are important mathematical operations for applications that rely on multidimensional data. The tensor-vector multiplication (TVM) is the most memory-bound tensor contraction in this class of operations. This article proposes an open-source TVM algorithm which is much simpler and efficient than previous approaches, making it suitable for integration in the most popular BLAS libraries available today. Our algorithm has been written from scratch and features unit-stride memory accesses, cache awareness, mode obliviousness, full vectorization and multi-threading as well as NUMA awareness for non-hierarchically stored dense tensors. Numerical experiments are carried out on tensors up to order 10 and various compilers and hardware architectures equipped with traditional DDR and high bandwidth memory (HBM). For large tensors the average performance of the TVM ranges between 62% and 76% of the theoretical bandwidth for NUMA systems with DDR memory and remains independent of the contraction mode. On NUMA systems with HBM the TVM exhibits some mode dependency but manages to reach performance figures close to peak values. Finally, the higher-order power method is benchmarked with the proposed TVM kernel and delivers on average between 58% and 69% of the theoretical bandwidth for large tensors.
Keyword:
Tensors
Kernel
Libraries
Bandwidth
Virtual machine monitors
Layout
Benchmark testing
Parallel algorithms
shared memory
tensor computations
high bandwidth memory
NUMA

期刊

IEEE Transactions on Parallel and Distributed Systems 封面图
IEEE Transactions on Parallel and Distributed Systems
IF:
6
论文数:
5.2K
被引数:
1.1W

机构

U
universitat politecnica de catalunya
学者数:
1.9W
论文数: 1.6W
被引数: 17
引用论文

引用论文

Research progress of efficient green perovskite light emitting diodes
err2019-01-01
err0
errOAAI
errZi-Han Qu; Ze-Ma Chu; Xing-Wang Zhang; Jing-Bi You
err分享
err收藏
Challenges of Memory Management on Modern NUMA Systems
err2015-11-23
err31
PREAI
errGaud, Fabien; Lepers, Baptiste; Funston, Justin; Dashti, Mohammad; Fedorova, Alexandra; Quema, Vivien; Lachaize, Renaud; Roth, Mark
err分享
err收藏
err分享
err收藏
Tensor Decompositions and Applications张量分解及其应用
err2009-08-05
err7.5K
PREAI
errKolda, Tamara G.; Bader, Brett W.
err分享
err收藏
学者 查看更多内容