arrow
Return

Fault-Aware Runtime Strategies for High-Performance Computing

delete2009-04-01
delete17
delete
OA
AI
Y
Yawei Li *
Z
Zhiling Lan
X
Xian‐He Sun
DOI:10.1109/TPDS.2008.128delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
As the scale of parallel systems continues to grow, fault management of these systems is becoming a critical challenge. While existing research mainly focuses on developing or improving fault tolerance techniques, a number of key issues remain open. In this paper, we propose runtime strategies for spare node allocation and job rescheduling in response to failure prediction. These strategies, together with failure prediction and fault tolerance techniques, construct a runtime system called Fault-Aware Runtime System (FARS). In particular, we propose a 0-1 knapsack model and demonstrate its flexibility and effectiveness for reallocating running jobs to avoid failures. Experiments, by means of synthetic data and real traces from production systems, show that FARS has the potential to significantly improve system productivity (i.e., performance and reliability).
Keywords:
High-performance computing
runtime strategies
fault tolerance
performance
reliability
0-1 knapsack
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

I
Illinois Institute of Technology
Scholars:
3.8K
Papers: 3.9K
Citations: 4.2K