arrow
Return

A Survey on Automatic Parameter Tuning for Big Data Processing Systems

delete2020-04-26
delete62
delete
OA
AI
H
Herodotos Herodotou *
Y
Yuxing Chen
J
Jiaheng Lu
DOI:10.1145/3381027delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Big data processing systems (e.g., I Iadoop, Spark, Storm) contain a vast number of configuration parameters controlling parallelism, I/O behavior, memory settings, and compression. Improper parameter settings can cause significant performance degradation and stability issues. However, regular users and even expert administrators grapple with understanding and tuning them to achieve good performance. We investigate existing approaches on parameter tuning for both batch and stream data processing systems and classify them into six categories: rule-based, cost modeling, simulation-based, experiment-driven, machine learning, and adaptive tuning. We summarize the pros and cons of each approach and raise some open research problems for automatic parameter tuning.
Keywords:
Parameter tuning
self-tuning
MapReduce
Spark
Storm
stream
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

ACM Computing Surveys cover
ACM Computing Surveys
IF:
28
Papers:
2.4K
Citations:
3.5W

Organization

U
university of helsinki
Scholars:
4.1W
Papers: 3.6W
Citations: 51
C
Cyprus University of Technology
Scholars:
2.1K
Papers: 2.1K
Citations: 2.6K