返回
Towards Generic Failure-Prediction Models in Large-Scale Distributed Computing Systems
DOI:10.3390/electronics14173386.png)
摘要
En 中文
分布式计算(DC)系统日益增长的复杂性需要先进的故障预测模型来提升可靠性和效率。本研究提出了一种全面的方法论,用于开发能够实现跨层和跨平台故障预测的通用机器学习(ML)模型,且无需进行平台特定的再训练。我们使用来自故障跟踪档案(FTA)的Grid5000故障数据集,探索了线性回归、逻辑回归、随机森林和XGBoost等方法,以预测三个关键指标:故障间隔时间(TBF)、恢复/修复时间(TTR)和故障节点识别(FNI)。我们的方法包括广泛的探索性数据分析(EDA)、故障模式的统计检验以及模型在集群、站点和系统层面的评估。结果表明,XGBoost始终优于其他模型,在TBF和FNI上实现了接近完美的100%准确率,并在多样化的DC环境中表现出良好的泛化能力。此外,我们引入了一种分层DC架构,整合了这些故障预测模型。以用例的形式,我们还展示了服务提供商如何利用这些预测模型来平衡服务可靠性和成本。
Keyword:
failure prediction
machine learning
distributed computing
XGBoost
cross-platform modeling
期刊
IF:
2.6
论文数:
9.9K
被引数:
4.7W
机构
引用论文
Predictive modelling of MapReduce job performance in cloud environments using machine learning techniques使用机器学习技术对云环境中的MapReduce作业性能进行预测建模
JOURNAL OF BIG DATA
IF6.4
Gaykar, R.S.; Khanaa, V.; Joshi, S.D. A hybrid supervised learning approach for detection and mitigation of job failure with virtual machines in distributed environments. Ing. Des Syst. D’Inf. 2022, 27, 621. [Google Scholar] [CrossRef]Gaykar, R.S.; Khanaa, V.; Joshi, S.D. 一种用于分布式环境中虚拟机作业失败检测与缓解的混合监督学习方法。Ing. Des Syst. D’Inf. 2022, 27, 621. [Google Scholar] [CrossRef]
Hochenbaum, J.; Vallis, O.S.; Kejariwal, A. Automatic anomaly detection in the cloud via statistical learning. arXiv 2017, arXiv:1704.07706. [Google Scholar] [CrossRef]Hochenbaum, J.; Vallis, O.S.; Kejariwal, A. 通过统计学习方法实现云环境中的自动异常检测。arXiv 2017, arXiv:1704.07706. [Google Scholar] [CrossRef]

