arrow
Return

Processor allocation and checkpoint interval selection in cluster computing systems

delete2001-11-01
delete49
PRE
AI
J
James S. Plank *
M
Michael G. Thomason
DOI:10.1006/jpdc.2001.1757delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Performance prediction of checkpointing systems in the presence of failures is a well-studied research area. While the literature abounds with performance models of checkpointing systems, none addresses the issue of selecting runtime parameters other than the optimal checkpointing interval. In particular, the issue of processor allocation is typically ignored. In this paper, we present a performance model for long-running parallel computations that execute with checkpointing enabled. We then discuss how it is relevant to today's parallel computing environments and software, and present case studies of using the model to select runtime parameters. (C) 2001 Academic Press.
Keywords:
checkpointing
performance prediction
parameter selection
parallel computation
Markov chain
exponential failure
exponential repair

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

No organization information available