arrow
返回

A new hybrid GPU-CPU sparse LDLT factorization algorithm with GPU and CPU factorizing concurrently

delete2024-07-01
delete1
PRE
AI
Y
Yunmou Liu
杜
杜辉 (Hui Du)
Z
Zhuogen Li
P
Pu Chen *
DOI:10.1016/j.jocs.2024.102312delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This paper proposes a new task assignment scheme for sparse LDLT factorization on a hybrid GPU-CPU platform, and correspondingly an efficient supernodal algorithm, GPU Node and CPU Pipelining (GNCP). GNCP assigns a large number of weakly coupled tasks to GPU and CPU respectively so that GPU and CPU can execute different factorization tasks concurrently without explicit synchronization, which has not been achieved by previous researches. More precisely, based on the number of CPU threads, the global memory of GPU as well the size of the matrix, we introduce the concept of the truncation level in the partition tree built by the multi-level graph partitioning. The sum of FLOP associated with decoupled nodes in the truncation level accounts for about 65% of the whole computations in tested numerical examples. GNCP assigns the factorization tasks of each node in the truncation level to GPU. The data of each node in the truncation level are copied to the global memory of GPU and factorized entirely and independently by GPU one-by-one. Once factorization of any node in the truncation level has been completed, the cross-factorizations of this node to its ancestors and the following related operations are decoupled from other GPU tasks. These tasks are assigned to CPU threads with a reentrant pipelining strategy, accounting for about 35% of the whole calculations. Using this novel design, CPU and GPU are working concurrently and a large part of factorization is performed in overlapped time. Numerical tests for matrices from SuiteSparse matrix collection show higher performance than the well-known hybrid GPU-CPU solver CHOLMOD.
Keyword:
Sparse LDL T factorization
GPU algorithm
Hybrid GPU-CPU algorithm
Parallel algorithm
Partition tree

期刊

Nature Computational Science 封面图
Nature Computational Science
IF:
18.3
论文数:
3.1K
被引数:
4.0K

机构

P
peking university
学者数:
11.9W
论文数: 8.7W
被引数: 146
引用论文

引用论文

Next-Generation Transparent Conducting Oxides for Photovoltaic Cells: an Overview
err2011-03-21
err0
PREAI
errDavid Ginley; Tim Coutts; John Perkins; David Young; Xiaonan Li; Phil Parilla
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容