arrow
Return

A Fault-Tolerant Framework for Asynchronous Iterative Computations in Cloud Environments

delete2018-08-01
delete10
PRE
AI
Z
Zhigang Wang *
L
Lixin Gao
Y
Yu Gu
鲍玉斌 cover
鲍玉斌 (Yubin Bao)
G
Ge Yu
DOI:10.1109/TPDS.2018.2808519delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Most graph algorithms are iterative in nature. They can be processed by distributed systems in memory in an efficient asynchronous manner. However, it is challenging to recover from failures in such systems. This is because traditional checkpoint fault-tolerant frameworks incur expensive barrier costs that usually offset the gains brought by asynchronous computations. Worse, surviving data are rolled back, leading to costly re-computations. This paper first proposes to leverage surviving data for failure recovery in an asynchronous system. Our framework guarantees the correctness of algorithms and avoids rolling back surviving data. Additionally, a novel asynchronous checkpointing solution is introduced to accelerate recovery at the price of nearly zero overheads. Some optimization strategies like message pruning, non-blocking recovery and load balancing are also designed to further boost the performance. We have conducted extensive experiments to show the effectiveness of our proposals using real-world graphs.
Keywords:
Fault-tolerance
asynchronous model
iterative graph algorithm
distributed memory-based systems
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

U
university of massachusetts system
Scholars:
3.8W
Papers: 3.5W
Citations: 42
N
northeastern university - china
Scholars:
3.1W
Papers: 2.7W
Citations: 37