arrow
Return

Reliability-oriented resource management for High-Performance Computing

delete2023-09-01
delete2
delete
OA
AI
G
Giuseppe Massari *
M
Miriam Peta
A
Alessandro Campi
F
Federico Reghenzani
F
Federico Terraneo
G
Giovanni Agosta
W
William Fornaciari
S
Sebastian Ciesielski
M
Michał Kulczewski
W
Wojciech Piątek
DOI:10.1016/j.suscom.2023.100873delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Reliability is an increasingly pressing issue for High-Performance Computing systems, as failures are a threat to large-scale applications, for which an even single run may incur significant energy and billing costs. Currently, application developers need to address reliability explicitly, by integrating application-specific checkpoint/restore mechanisms. However, the application alone cannot exploit system knowledge, which is not the case for system-wide resource management systems. In this paper, we propose a reliability-oriented policy that can increase significantly component reliability by combining checkpoint/restore mechanisms exploitation and proactive resource management policies.
Keywords:
Reliability
HPC
Distributed systems
Resource management
Software simulators
Thermal management
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

S
Sustainable Computing-Informatics and Systems
IF:
5.7
Papers:
966
Citations:
2.9K

Organization

P
Polish Academy of Sciences
Scholars:
3.0W
Papers: 3.1W
Citations: 3.1W
P
Polytechnic University of Milan
Scholars:
2.0W
Papers: 1.8W
Citations: 24