arrow
Return

Preparing MPICH for exascale

delete2025-01-09
delete0
PRE
AI
Y
Yanfei Guo *
K
Ken Raffenetti
H
Hui Zhou
M
Min Si
A
Abdelhalim Amer
S
Shintaro Iwasaki
S
Sangmin Seo
G
Giuseppe Congiu
R
Robert Latham
L
Lena Oden
T
Thomas Gillis
R
Rohit Zambre
K
Kaiming Ouyang
C
Charles J Archer
W
Wesley Bland
J
Jithin Jose
S
Sayantan Sur
H
Hajime Fujita
D
Dmitry Durnov
M
Michael Chuvelev
S
Sagar Thapaliya
T
Taru Doodi
M
Maria Garazan
S
Steve Oyanagi
M
Marc Snir
R
Rajeev Thakur
DOI:10.1177/10943420241311608delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The advent of exascale supercomputers heralds a new era of scientific discovery, yet it introduces significant architectural challenges that must be overcome for MPI applications to fully exploit its potential. Among these challenges is the adoption of heterogeneous architectures, particularly the integration of GPUs to accelerate computation. Additionally, the complexity of multithreaded programming models has also become a critical factor in achieving performance at scale. The efficient utilization of hardware acceleration for communication, provided by modern NICs, is also essential for achieving low latency and high throughput communication in such complex systems. In response to these challenges, the MPICH library, a high-performance and widely used Message Passing Interface (MPI) implementation, has undergone significant enhancements. This paper presents four major contributions that prepare MPICH for the exascale transition. First, we describe a lightweight communication stack that leverages the advanced features of modern NICs to maximize hardware acceleration. Second, our work showcases a highly scalable multithreaded communication model that addresses the complexities of concurrent environments. Third, we introduce GPU-aware communication capabilities that optimize data movement in GPU-integrated systems. Finally, we present a new datatype engine aimed at accelerating the use of MPI derived datatypes on GPUs. These improvements in the MPICH library not only address the immediate needs of exascale computing architectures but also set a foundation for exploiting future innovations in high-performance computing. By embracing these new designs and approaches, MPICH-derived libraries from HPE Cray and Intel were able to achieve real exascale performance on OLCF Frontier and ALCF Aurora respectively.
Keywords:
Message passing interface
MPI
HPC communication
HPC network
exascale MPI

Journal

International Journal of High Performance Computing Applications cover
International Journal of High Performance Computing Applications
IF:
2.5
Papers:
1.1K
Citations:
1.3K

Organization

I
intel usa
Scholars:
736
Papers: 548
Citations: 1
A
Argonne National Laboratory
Scholars:
1.1W
Papers: 9.2K
Citations: 3.8W
U
University of Illinois Urbana-Champaign
Scholars:
2.4W
Papers: 2.0W
Citations: 35
U
united states department of energy (doe)
Scholars:
11.3W
Papers: 9.6W
Citations: 246
University of California System cover
University of California System
Scholars:
37.5W
Papers: 33.7W
Citations: 6.6K
I
Intel Corporation
Scholars:
2.7K
Papers: 2.0K
Citations: 6
researcher View more organizations