arrow
Return

Quantifying event correlations for proactive failure management in networked computing systems

delete2010-11-01
delete33
PRE
AI
S
Song Fu *
C
Chengzhong Xu
DOI:10.1016/j.jpdc.2010.06.010delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Networked computing systems continue to grow in scale and in the complexity of their components and interactions. Component failures become norms instead of exceptions in these environments. Moreover, failure events exhibit strong correlations in the time and space domains. In this paper, we develop a spherical covariance model with an adjustable timescale parameter to quantify the temporal correlation and a stochastic model to characterize spatial correlation. The models are further extended to take into account the information of application allocation to discover more correlations among failure instances. We cluster failure events based on their correlations and predict their future occurrences. Experimental results on a production coalition system, the Wayne State Computational Grid, show the offline and online predictions made by our predicting system can forecast 72.7-85.3% of the failure occurrences and capture failure correlations in a cluster coalition environment. (c) 2010 Elsevier Inc. All rights reserved.
Keywords:
Failure characterization
Temporal correlation
Spatial correlation
System availability
Networked computing systems
Autonomic management
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

U
university of north texas denton
Scholars:
4.4K
Papers: 3.9K
Citations: 10
U
University of North Texas System
Scholars:
8.0K
Papers: 7.7K
Citations: 178