返回
Asynchronous recovery without using vector timestamps
DOI:10.1016/S0743-7315(02)00005-9.png)
摘要
En 中文
A checkpoint of a process involved in a distributed computation is said to be useful if it is part of a consistent global checkpoint. In this paper, we present a quasi-synchronous checkpointing algorithm that makes every checkpoint useful. We also present an efficient asynchronous recovery algorithm based on the checkpointing algorithm. The checkpointing algorithm allows the processes to take checkpoints asynchronously and also forces the processes to take additional checkpoints in order to make every checkpoint useful. The recovery algorithm can handle concurrent failure of multiple processes. The recovery algorithm has no domino effect and a failed process needs only to roll back to its latest checkpoint and request the other processes to roll back to a consistent checkpoint. Messages are only selectively logged to cope with various types of message abnormalities that arise due to rollback and hence results in low message logging overhead. Unlike some existing algorithms, our algorithm does not use vector timestamps for tracking dependency between checkpoints and hence results in low message overhead during failure-free operation. Moreover, a process can asynchronously decide garbage checkpoints and delete them from the stable storage-garbage checkpoints are the checkpoints that are no longer required for the purpose of recovery. (C) 2002 Elsevier Science (USA). All rights reserved.
Keyword:
distributed checkpointing
quasi-synchronous checkpointing
communication-induced check-pointing
failure-recovery
fault-tolerance
rollback-recovery
vector timestamps
multiple failures
asynchronous recovery
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
4
论文数:
3.8K
被引数:
4.8K
机构
暂无机构信息

