arrow
Return

Performance evaluation of hybrid programming patterns for large CPU/GPU heterogeneous clusters

delete2012-06-01
delete23
PRE
AI
F
Fengshun Lu *
J
Junqiang Song
F
Fukang Yin
X
Xiaoqian Zhu
DOI:10.1016/j.cpc.2012.01.019delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The CPU/GPU heterogeneous clusters are important platforms for high performance computing applications. However, there are many challenges for efficiently performing the scientific and engineering legacy code on these heterogeneous systems. In this paper, we endeavor to address the programming-model issue by combining the existing models (i.e., MPI, OpenMP and CUDA). First, two hybrid programming patterns are presented, namely the MPI + CUDA and MPI + OpenMP/CUDA. Second, three kernels (i.e., EP, CG and MG) of the NAS parallel benchmarks (NPBs), which are abstracted from many legacy computational fluid dynamics applications, are implemented with the above two patterns. Third, these hybrid implementations are executed on the TianHe-1A supercomputer, and the corresponding experimental results show that significant performance improvement can be achieved with the above patterns. Finally, a detailed performance analysis about the two hybrid patterns is performed and some guidelines for porting the legacy code onto large-scale heterogeneous CPU/GPU clusters are also given. (C) 2012 Elsevier B.V. All rights reserved.
Keywords:
MPI
CUDA
OpenMP
GPU cluster
NPB
Performance evaluation
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Computer Physics Communications cover
Computer Physics Communications
IF:
3.4
Papers:
1.2W
Citations:
3.7W

Organization

N
national university of defense technology - china
Scholars:
1.8W
Papers: 1.4W
Citations: 9