arrow
Return

Multi-GPU work sharing in a task-based dataflow programming model

delete2024-07-01
delete1
delete
OA
AI
J
Joseph John *
J
Josh Milthorpe
T
Thomas Hérault
G
George Bosilca
DOI:10.1016/j.future.2024.03.017delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Today, multi-GPU computing nodes are the mainstay of most high-performance computing systems. Despite significant progress in programmability, building an application that efficiently utilizes all the GPUs in a computing node is still a significant challenge, especially using the existing shared memory and messagepassing paradigms. In this aspect, the task -based dataflow programming model has emerged as an alternative for multi-GPU computing nodes. Most task -based dataflow runtimes have dynamic task mapping, where tasks are mapped to different GPUs based on the current load, but once the mapping has been established, there is no rebalancing of tasks even if an imbalance is detected. In this paper, we examine how automatic dynamic work sharing between GPUs within a compute node can improve the performance of an application through better workload distribution. We demonstrate performance improvement through dynamic work sharing using a Block -Sparse GEneral Matrix Multiplication (BSpGEMM) benchmark. Although we use PaRSEC, a task -based dataflow runtime, as the vehicle for this research, the ideas discussed here are transferable to any task -based dataflow runtime.
Keywords:
Tasks
Runtime
Work sharing
PaRSEC
GPU
Load balancing
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
Papers:
6.8K
Citations:
2.3W

Organization

A
Australian National University
Scholars:
2.1W
Papers: 2.3W
Citations: 3.9W
University of Tennessee System cover
University of Tennessee System
Scholars:
2.9W
Papers: 2.6W
Citations: 115