arrow
返回

PaScaL_TDMA 2.0: A multi-GPU-based library for solving massive tridiagonal systems

delete2023-09-01
delete1
PRE
AI
M
Mingyu Yang
J
Ji-Hoon Kang
K
Ki-Ha Kim
O
Oh‐Kyoung Kwon
J
Jung‐Il Choi *
DOI:10.1016/j.cpc.2023.108785delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
We introduce an updated library, PaScaL_TDMA 2.0, which was originally designed for the efficient computation of batched tridiagonal systems and is now capable of exploiting multi-GPU environments. The library extends its functionality to include GPU support and minimizes CPU-GPU data transfer by utilizing the device-resident memory while retaining the original CPU-based capabilities. The library employs pipeline copying with shared memory for low-latency memory access and incorporates CUDA-aware MPI for efficient multi-GPU communication. Our GPU implementation demonstrated outstanding computational performance compared to the original CPU implementation while consuming much less energy. In summary, this updated version presents a time-efficient and energy-saving approach for solving batched tridiagonal systems on modern computing platforms, including both GPU and CPU.
Keyword:
CUDA
GPU computing
Multi-GPU
Tridiagonal matrix systems

期刊

Computer Physics Communications 封面图
Computer Physics Communications
IF:
3.4
论文数:
1.2W
被引数:
3.7W

机构

K
korea institute of science & technology information (kisti)
学者数:
1.1K
论文数: 998
被引数: 0
Y
Yonsei University
学者数:
4.8W
论文数: 4.6W
被引数: 5.2W