arrow
返回

Data-Efficient Performance Modeling for Configurable Big Data Frameworks by Reducing Information Overlap Between Training Examples

delete2022-11-01
delete2
PRE
AI
Z
Zhiqiang Liu
X
Xuanhua Shi
金海 (Hai Jin) *
DOI:10.1016/j.bdr.2022.100358delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
To support the various analysis application of big data, big data processing frameworks are designed to be highly configurable. However, for common users, it is difficult to tailor the configurable frameworks to achieve optimal performance for every application. Recently, many automatic tuning methods are proposed to configure these frameworks. In detail, these methods firstly build a performance prediction model through sampling configurations randomly and measuring the corresponding performance. Then, they conduct heuristic search in the configuration space based on the performance prediction model. For most frameworks, it is too expensive to build the performance model since it needs to measure the performance of large amounts of configurations, which cause too much overhead on data collection. In this paper, we propose a novel data-efficient method to build the performance model with little impact on prediction accuracy. Compared to the traditional methods, the proposed method can reduce the overhead of data collection because it can train the performance model with much less training examples. Specifically, the proposed method can actively sample the important examples according to the dynamic requirement of the performance model during the iterative model updating. Hence, it can make full use of the collected informative data and train the performance model with much less training examples. To sample the important training examples, we employ several virtual performance model to estimate the importance of all candidate configurations efficiently. Experimental results show that our method needs less training examples than traditional methods with little impact on prediction accuracy.(c) 2022 Elsevier Inc. All rights reserved.
Keyword:
Big data framework
Highly configurable software
Performance model
Active sampling

期刊

Big Data Research 封面图
Big Data Research
IF:
4.2
论文数:
406
被引数:
1.1K

机构

暂无机构信息
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
RFHOC: A Random-Forest Approach to Auto-Tuning Hadoop's Configuration
err2016-05-01
err75
errOAAI
errBei, Zhendong; Yu, Zhibin; Zhang, Huiling; Xiong, Wen; Xu, Chengzhong; Eeckhout, Lieven; Feng, Shengzhong
err分享
err收藏
Soft computing techniques for big data and cloud computing
err2020-03-06
err6
errOAAI
errGupta, B. B.; Agrawal, Dharma P.; Yamaguchi, Shingo; Sheng, Michael
err分享
err收藏
Ligand Strain and Its Conformational Complexity Is a Major Factor in the Binding of Cyclic Dinucleotides to STING Protein
err2021-03-24
err0
errOAAI
errMiroslav Smola; Ondrej Gutten; Milan Dejmek; Milan Kožíšek; Thomas Evangelidis; Zahra Aliakbar Tehrani; Barbora Novotná; Radim Nencka; Gabriel Birkuš; Lubomír Rulíšek; Evzen Boura
err分享
err收藏
err分享
err收藏
没有更多内容