Return
Reliability-oriented resource management for High-Performance Computing
DOI:10.1016/j.suscom.2023.100873.png)
Abstract
En 中文
Reliability is an increasingly pressing issue for High-Performance Computing systems, as failures are a threat to large-scale applications, for which an even single run may incur significant energy and billing costs. Currently, application developers need to address reliability explicitly, by integrating application-specific checkpoint/restore mechanisms. However, the application alone cannot exploit system knowledge, which is not the case for system-wide resource management systems. In this paper, we propose a reliability-oriented policy that can increase significantly component reliability by combining checkpoint/restore mechanisms exploitation and proactive resource management policies.
Keywords:
Reliability
HPC
Distributed systems
Resource management
Software simulators
Thermal management
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
S
IF:
5.7
Papers:
966
Citations:
2.9K

