arrow
Return

A utilization model for optimization of checkpoint intervals in distributed stream processing systems

delete2020-09-01
delete16
delete
OA
AI
S
Sachini Jayasekara *
A
Aaron Harwood
S
Shanika Karunasekera
DOI:10.1016/j.future.2020.04.019delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
State-of-the-art distributed stream processing systems such as Apache Flink and Storm have recently included checkpointing to provide fault-tolerance for stateful applications. This is a necessary eventuality as these systems head into the Exascale regime, and is evidently more efficient than replication as state size grows. However current systems use a nominal value for the checkpoint interval, indicative of assuming roughly 1 failure every 19 days, that does not take into account the salient aspects of the checkpoint process, nor the system scale, which can readily lead to inefficient system operation. To address this shortcoming, we provide a rigorous derivation of utilization - the fraction of total time available for the system to do useful work - that incorporates checkpoint interval, failure rate, checkpoint cost, failure detection and restart cost, depth of the system topology and message delay. Our model yields an elegant expression for utilization and provides an optimal checkpoint interval given these parameters, interestingly showing it to be dependent only on checkpoint cost and failure rate. We confirm the accuracy and efficacy of the model through simulations and experiments with Apache Flink. Observations of the simulations validate our theoretical model and demonstrate that utilization can be improved using the derived optimal checkpoint interval. Moreover, experimental results with Apache Flink show that we can obtain improvements in system utilization for every case we tested, especially as the system size increases. Our model provides a solid theoretical basis for the analysis and optimization of more elaborate checkpointing approaches. (C) 2020 Elsevier B.V. All rights reserved.
Keywords:
Fault tolerance
Stream processing
Distributed systems
Checkpoint
Optimization
Modeling
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
Papers:
6.8K
Citations:
2.3W

Organization

U
university of melbourne
Scholars:
5.7W
Papers: 5.4W
Citations: 69