arrow
返回

Performance Portable Batched Sparse Linear Solvers

delete2023-05-01
delete1
delete
OA
AI
K
Kim Liegeois *
S
Sivasankaran Rajamanickam
L
Luc Berger‐Vergiat
DOI:10.1109/TPDS.2023.3249110delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Solving large number of small linear systems is increasingly becoming a bottleneck in computational science applications. While dense linear solvers for such systems have been studied before, batched sparse linear solvers are just starting to emerge. In this paper, we discuss algorithms for solving batched sparse linear systems and their implementation in the Kokkos Kernels library. The new algorithms are performance portable and map well to the hierarchical parallelism available in modern accelerator architectures. The sparse matrix vector product (SPMV) kernel is the main performance bottleneck of the Krylov solvers we implement in this work. The implementation of the batched SPMV and its performance are therefore discussed thoroughly in this paper. The implemented kernels are tested on different Central Processing Unit (CPU) and Graphic Processing Unit (GPU) architectures. We also develop batched Conjugate Gradient (CG) and batched Generalized Minimum Residual (GMRES) solvers using the batched SPMV. Our proposed solver was able to solve 20,000 sparse linear systems on V100 GPUs with a mean speedup of 76x and 924x compared to using a parallel sparse solver with a block diagonal system with all the small linear systems, and compared to solving the small systems one at a time, respectively. We see mean speedup of 0.51 compared to dense batched solver of cuSOLVER on V100, while using lot less memory. Thorough performance evaluation on three different architectures and analysis of the performance are presented.
Keyword:
Linear systems
Kernel
Graphics processing units
Tensors
Sparse matrices
Libraries
Instruction sets
Batch sparse solvers
batch BLAS
kokkos kernels
performance portable

期刊

IEEE Transactions on Parallel and Distributed Systems 封面图
IEEE Transactions on Parallel and Distributed Systems
IF:
6
论文数:
5.2K
被引数:
1.1W

机构

U
united states department of energy (doe)
学者数:
11.3W
论文数: 9.6W
被引数: 246
引用论文

引用论文

Tensor Decompositions and Applications张量分解及其应用
err2009-08-05
err7.5K
PREAI
errKolda, Tamara G.; Bader, Brett W.
err分享
err收藏
err分享
err收藏
Kokkos 3: Programming Model Extensions for the Exascale Era
err2022-04-01
err165
errOAAI
errTrott, Christian R.; Lebrun-Grandie, Damien; Arndt, Daniel; Ciesko, Jan; Dang, Vinh; Ellingwood, Nathan; Gayatri, Rahulkumar; Harvey, Evan; Hollman, Daisy S.; Ibanez, Dan; Liber, Nevin; Madsen, Jonathan; Miles, Jeff; Poliakoff, David; Powell, Amy; Rajamanickam, Sivasankaran; Simberg, Mikael; Sunderland, Dan; Turcksin, Bruno; Wilke, Jeremiah
err分享
err收藏
Observation of polarization effects in Λc+ semileptonic decay
err1994-05-01
err0
PREAI
errH. Albrecht; H. Ehrlichmann; T. Hamacher; R.P. Hofmann; T. Kirchhoff; R. Mankel; A. Nau; S. Nowak; H. Schröder; H.D. Schulz; M. Walter; R. Wurth; C. Hast; H. Kapitza; H. Kolanoski; A. Kosche; A. Lange; A. Lindner; M. Schieber; T. Siegmund; B. Spaan; H. Thurn; D. Töpfer; D. Wegener; P. Eckstein; K.R. Schubert; R. Schwierz; R. Waldi; K. Reim; H. Wegener; R. Eckmann; H. Kuipers; O. Mai; R. Mundt; T. Oest; R. Reiner; W. Schnidt-Parzefall; J. Stiewe; S. Werner; K. Ehret; W. Hofmann; A. Hüpper; S. Khan; K.T. Knöpfle; M. Seeger; J. Spengler; P. Krieger; D.B. MacFarlane; J.D. Prentice; P.R.B. Saull; K. Tzamariudaki; R.G. Van de Water; T.-S. Yoon; C. Frankl; D. Reβing; M. Schmidtler; M. Schneider; S. Weseler; G. Kernel; P. Križan; E. Križnič; T. Podobnik; T. Živko; V. Balagura; I. Belyaev; S. Chechelnitsky; M. Danilov; A. Droutskoy; Yu. Gershtein; A. Golutvin; I. Korolko; G. Kostina; D. Litvintsev; V. Lubimov; P. Pakhlov; S. Semenov; A. Snizhko; I. Tichomirov; Yu. Zaitsev
err分享
err收藏
学者 查看更多内容