arrow
返回

Process fault tolerance: Semantics, design and applications for high performance computing

delete2005-11-01
delete27
PRE
AI
G
Graham E. Fagg
E
Edgar Gabriel
Z
Zizhong Chen
T
Thara Angskun
B
Bosillca, G
J
Jelena Pješivac–Grbović
D
Dongarra, JJ
DOI:10.1177/1094342005056137delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
With increasing numbers of processors on current machines, the probability for node or link failures is also increasing. Therefore, application-level fault tolerance is becoming more of an important issue for both end-users and the institutions running the machines. In this paper we present the semantics of a fault-tolerant version of the message passing interface (MPI), the de-facto standard for communication in scientific applications, which gives applications the possibility to recover from a node or link error and continue execution in a well-defined way. We present the architecture of fault-tolerant MPI, an implementation of MPI using the semantics presented above as well as benchmark results with various applications. An example of a fault-tolerant parallel equation solver, performance results as well as the time for recovering from a process failure are furthermore detailed.
Keyword:
parallel computing
fault tolerance
MPI and message passing
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

International Journal of High Performance Computing Applications 封面图
International Journal of High Performance Computing Applications
IF:
2.5
论文数:
1.1K
被引数:
1.3K

机构

暂无机构信息
引用论文

引用论文

Follow-up Assessment of Standing Mobility Device Users站立移动设备使用者的随访评估
err2010-10-22
err0
PREAI
errRobert B. Dunn; James S. Walter; Yuvone Lucero; Frances Weaver; Edwin Langbein; Linda Fehr; Paul Johnson; Lisa Riedy
err分享
err收藏