arrow
返回

Efficient heterogeneous execution on large multicore and accelerator platforms: Case study using a block tridiagonal solver

delete2013-12-01
delete4
PRE
AI
A
Alfred J. Park
K
Kalyan S. Perumalla *
DOI:10.1016/j.jpdc.2013.07.012delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The algorithmic and implementation principles are explored in gainfully exploiting GPO accelerators in conjunction with multicore processors on high-end systems with large numbers of compute nodes, and evaluated in an implementation of a scalable block tridiagonal solver. The accelerator of each compute node is exploited in combination with multicore processors of that node in performing block-level linear algebra operations in the overall, distributed solver algorithm. Optimizations incorporated include: (1) an efficient memory mapping and synchronization interface to Minimize data movement, (2) multi-process sharing of the accelerator within a node to obtain balanced load with multicore processors', and (3) an automatic memory management system to efficiently utilize accelerator memory when sub-matrices spill over the limits of device memory. Results are reported from our novel implementation that uses MAGMA and CUBLAS accelerator software systems simultaneously with ACML (2013) [2] for multithreaded execution on processors. Overall, using 940 nVidia Testa X2090 accelerators and 15,040 cores, the best heterogeneous execution delivers a 10.9-fold reduction in run time relative to an already efficient parallel multicore-only baseline implementation that is highly optimized with intra-node and inter-node concurrency and computation-communication overlap. Detailed quantitative results are presented to explain all critical runtime components contributing to hybrid performance. (C) 2013 Elsevier Inc. All rights reserved.
Keyword:
Tridiagonal solver
Linear algebra
GPU
Accelerator
Heterogeneous execution
Memory management

期刊

Journal of Parallel and Distributed Computing 封面图
Journal of Parallel and Distributed Computing
IF:
4
论文数:
3.8K
被引数:
4.8K

机构

U
united states department of energy (doe)
学者数:
11.3W
论文数: 9.6W
被引数: 246
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
err分享
err收藏
err分享
err收藏
GPU-based parallel algorithms for sparse nonlinear systems
err2012-09-01
err15
PREAI
errGaliano, V.; Migallon, H.; Migallon, V.; Penades, J.
err分享
err收藏
Density functional theory calculation on many-cores hybrid central processing unit-graphic processing unit architectures
err2009-07-16
err87
errOAAI
errGenovese, Luigi; Ospici, Matthieu; Deutsch, Thierry; Mehaut, Jean-Francois; Neelov, Alexey; Goedecker, Stefan
err分享
err收藏
学者 查看更多内容