arrow
Return

Performance models for asynchronous data transfers on consumer Graphics Processing Units

delete2012-09-01
delete29
PRE
AI
J
Juan Gómez-Luna *
J
José María González-Linares
J
J.I. Benavides
N
Nicolás Guil
DOI:10.1016/j.jpdc.2011.07.011delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Graphics Processing Units (CPU) have impressively arisen as general-purpose coprocessors in high performance computing applications, since the launch of the Compute Unified Device Architecture (CUDA). However, they present an inherent performance bottleneck in the fact that communication between two separate address spaces (the main memory of the CPU and the memory of the CPU) is unavoidable. The CUDA Application Programming Interface (API) provides asynchronous transfers and streams, which permit a staged execution, as a way to overlap communication and computation. Nevertheless, a precise manner to estimate the possible improvement due to overlapping does not exist, neither a rule to determine the optimal number of stages or streams in which computation should be divided. In this work, we present a methodology that is applied to model the performance of asynchronous data transfers of CUDA streams on different CPU architectures. Thus, we illustrate this methodology by deriving expressions of performance for two different consumer graphic architectures belonging to the more recent generations. These models permit programmers to estimate the optimal number of streams in which the computation on the CPU should be broken up, in order to obtain the highest performance improvements. Finally, we have checked the suitability of our performance models with three applications based on codes from the CUDA Software Development Kit (SDK) with successful results. (C) 2011 Elsevier Inc. All rights reserved.
Keywords:
GPU
CUDA
Asynchronous transfers
Streams
Overlapping of communication and computation

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

U
universidad de malaga
Scholars:
1.2W
Papers: 9.2K
Citations: 6
U
universidad de cordoba
Scholars:
1.0W
Papers: 8.4K
Citations: 6
Cited Papers

Cited Papers