arrow
Return

Generating data transfers for distributed GPU parallel programs

delete2013-12-01
delete5
PRE
AI
F
Frédérique Silber-Chaussumier *
A
A. Müller
R
Rachid Habel
DOI:10.1016/j.jpdc.2013.07.022delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Nowadays, high performance applications exploit multiple level architectures, due to the presence of hardware accelerators like GPUs inside each computing node. Data transfers occur at two different levels: inside the computing node between the CPU and the accelerators and between computing nodes. We consider the case where the intra-node parallelism is handled with HMPP compiler directives and message-passing programming with MPI is used to program the inter-node communications. This way of programming on such an heterogeneous architecture is costly and error-prone. In this paper, we specifically demonstrate the transformation of HMPP programs designed to exploit a single computing node equipped with a GPU into an heterogeneous HMPP MPI exploiting multiple GPUs located on different computing nodes. The STEP tool focuses on generating communications combining both powerful static analyses and runtime execution to reduce the volume of communications. Our source-to-source transformation is implemented inside the PIPS workbench. We detail the generated source program of the Jacobi kernel and show that the execution times and speedups are encouraging. At last we give some directions for the improvement of the tool. (C) 2013 Elsevier Inc. All rights reserved.
Keywords:
Distributed memory
Data transfer
Source-to-source transformation
Parallel execution
Compiler directives
GPU

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

I
imt - institut mines-telecom
Scholars:
7.4K
Papers: 6.4K
Citations: 5