arrow
Return

Extending an asynchronous runtime system for high throughput applications: A case study

delete2022-05-01
delete0
delete
OA
AI
J
Joshua Suetterlein
J
Joseph Manzano *
A
Andrés Márquez
G
Guang R. Gao
DOI:10.1016/j.jpdc.2022.01.027delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Current supercomputers are mostly composed of vast numbers of nodes enhanced with accelerators (usually in the form of GPUs). However, having these heterogeneous designs in the forefront have exposed the software toolchains and application designers to the underlying complexities of an ever evolving hardware substrate. The need for a more dynamic view from the system software (i.e., compilers and runtimes) has become more apparent in these environments. Due to this, adaptive, fine grain runtime systems have seen a rise in popularity in the past decades. With low overhead and small tasks, these runtimes help to hide long latency operations by exploiting the massive concurrency presented in different application workflows. Such features allow the reduction of idle time (a result from ever deeper and complex memory hierarchies and memory types) with the execution of unrelated work across the machine. Of these runtimes, the Asynchronous Many Task (AMT) Runtimes are excellent exemplars as they can efficiently map onto hardware substrates and exhibit a high degree of latency hiding. Because of their latency tolerant characteristics, applications such as Graph Analytics and Big Data applications (which are latency sensitive) can use these runtimes very efficiently. Thanks to these characteristics, we present how a careful design can help to exploit the properties of an AMT when running high latency applications such as the ones encountered in the Big Data domain. In addition, when combined with introspection / adaptive capabilities, the runtime can further exploit optimization opportunities during its execution based on the ever changing state of the underlying hardware substrate. As a vehicle for this exploration, we use the Performance Open Community Runtime (P-OCR) to test all these concepts with Big Data workloads. (C)& nbsp;2022 Published by Elsevier Inc.
Keywords:
Big Data
Asynchronous runtime systems
Performance analysis
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

P
Pacific Northwest National Laboratory
Scholars:
9.0K
Papers: 6.3K
Citations: 14
U
united states department of energy (doe)
Scholars:
11.3W
Papers: 9.6W
Citations: 246