arrow
Return

Self-stabilizing algorithm for checkpointing in a distributed system

delete2007-07-01
delete3
PRE
AI
P
Partha Sarathi Mandal
K
Krishnendu Mukhopadhyaya *
DOI:10.1016/j.jpdc.2007.02.006delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
If the variables used for a checkpointing algorithm have data faults, the existing checkpointing and recovery algorithms may fail. In this paper, self-stabilizing data fault detecting and correcting, checkpointing, and recovery algorithms are proposed in a ring topology. The proposed data fault detection and correction algorithms can handle data faults; at most one per process, but in any number of processes. The proposed checkpointing algorithm can deal with concurrent multiple initiations of checkpointing and data faults. A process can recover from a fault, using the proposed recovery algorithm in spite of multiple data faults present in the system. All the proposed algorithms converge in O (n) steps, where n is the number of processes. The algorithm can be extended to work for general topologies too. (C) 2007 Elsevier Inc. All rights reserved.
Keywords:
data fault
process fault
checkpointing
rollback recovery
self-stabilization
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

No organization information available