Return
Quantifying rollback propagation in distributed checkpointing
DOI:10.1016/j.jpdc.2004.01.003.png)
Abstract
En 中文
This paper proposes a new classification of executions with checkpoints based on the amount of rollback during recovery. Specifically, an execution is k-rollback, if k indicates the maximal number of checkpoints that have to be rolled back. It is shown that coordinated checkpointing, SZPF, and ZPF are 1-rollback, while ZCF is (n - 1)-rollback, where n is the number of participants in an execution. A new class of executions, called d-bounded cycles (in short, d-BC). is introduced, and is shown to be ((n - 1) - d)-rollback (ZCF is a special case of d-BC for (d = 1). Finally, a protocol is presented whose executions are d-bounded cycles. A nice property of this protocol is that it does not impose any control information overhead on application messages, yet sends only a few control messages of its own. Moreover, the protocol maintains information that enables very efficient discovery of a recent recovery line that existed shortly before the failure. (C) 2004 Elsevier Inc. All rights reserved.
Keywords:
fault tolerance
checkpoint/restart
recovery lines
rollback propagation
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
4
Papers:
3.8K
Citations:
4.8K
Organization
No organization information available

