arrow
Return

Exploring the Design Tradeoffs for Extreme-Scale High-Performance Computing System Software

delete2016-04-01
delete10
delete
OA
AI
K
Ke Wang *
K
Kulkarni, Abhishek *
M
Michael Lang *
D
Dorian Arnold *
I
Ioan Raicu
DOI:10.1109/TPDS.2015.2430852delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Owing to the extreme parallelism and the high component failure rates of tomorrow's exascale, high-performance computing (HPC) system software will need to be scalable, failure-resistant, and adaptive for sustained system operation and full system utilizations. Many of the existing HPC system software are still designed around a centralized server paradigm and hence are susceptible to scaling issues and single points of failure. In this article, we explore the design tradeoffs for scalable system software at extreme scales. We propose a general system software taxonomy by deconstructing common HPC system software into their basic components. The taxonomy helps us reason about system software as follows: (1) it gives us a systematic way to architect scalable system software by decomposing them into their basic components; (2) it allows us to categorize system software based on the features of these components, and finally (3) it suggests the configuration space to consider for design evaluation via simulations or real implementations. Further, we evaluate different design choices of a representative system software, i.e. key-value store, through simulations up to millions of nodes. Finally, we show evaluation results of two distributed system software, Slurm++ (a distributed HPC resource manager) and MATRIX (a distributed task execution framework), both developed based on insights from this work. We envision that the results in this article help to lay the foundations of developing next-generation HPC system software for extreme scales.
Keywords:
Distributed systems
high-performance computing
key-value stores
simulation
systems and software
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

I
Illinois Institute of Technology
Scholars:
3.8K
Papers: 3.9K
Citations: 4.2K
I
indiana university system
Scholars:
4.0W
Papers: 3.5W
Citations: 38
I
Indiana University Bloomington
Scholars:
1.9W
Papers: 1.5W
Citations: 2.8W
U
united states department of energy (doe)
Scholars:
11.3W
Papers: 9.6W
Citations: 246
researcher View more organizations