arrow
返回

An intermediate data placement algorithm for load balancing in Spark computing environment

delete2018-01-01
delete50
PRE
AI
Z
Zhuo Tang *
Z
Zhang Xiang-shen
李肯立 封面图
李肯立 (Kenli Li)
李克勤 封面图
李克勤 (Keqin Li)
DOI:10.1016/j.future.2016.06.027delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Since MapReduce became an effective and popular programming framework for parallel data processing, key skew in intermediate data has become one of the important system performance bottlenecks. For solving the load imbalance of bucket containers in the shuffle process of the Spark computing framework, this paper proposes a splitting and combination algorithm for skew intermediate data blocks (SCID), which can improve the load balancing for various reduce tasks. Because the number of keys cannot be counted out until the input data are processed by map tasks, this paper provides a sampling algorithm based on reservoir sampling to detect the distribution of the keys in intermediate data. Contrasting with the original mechanism for bucket data loading, SOD sorts the data clusters of key/value tuples from each map task according to their sizes, and fills them into the relevant buckets orderly. A data cluster will be split once it exceeds the residual volume of the current bucket. After filling this bucket, the remainder cluster will be entered into the next iteration. Through this processing, the total size of data in each bucket is roughly scheduled equally. For each map task, each reduce task should fetch the intermediate results from a specific bucket, the quantity in all buckets for a map task will balance the load of the reduce tasks. We implement SCID in Spark 1.1.0 and evaluate its performance through three widely used benchmarks: Sort, Text Search, and Word Count. Experimental results show that our algorithms can not only achieve higher overall average balancing performance, but also reduce the execution time of a job with varying degrees of data skew. (C) 2016 Elsevier B.V. All rights reserved.
Keyword:
Data sampling
Data skew
Load balancing
MapReduce
Spark
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
论文数:
6.8K
被引数:
2.3W

机构

H
hunan university
学者数:
4.5W
论文数: 3.3W
被引数: 70
引用论文

引用论文

The Damaged Human Detrusor: Functional and Electron Microscopic Changes in Disease1
err1973-04-01
err0
PREAI
errM. E. MAYO; R. W. LLOYD-DAVIES; K. E. D. SHUTTLEWORTH; J. R. TIGHE
err分享
err收藏
Handling partitioning skew in MapReduce using LEEN
err2013-05-25
err49
errOAAI
errIbrahim, Shadi; Jin, Hai; Lu, Lu; He, Bingsheng; Antoniu, Gabriel; Wu, Song
err分享
err收藏
The grid workloads archive
err2008-07-01
err171
PREAI
errIosup, Alexandru; Li, Hui; Jan, Mathieu; Anoep, Shanny; Dumitrescu, Catalin; Wolters, Lex; Epema, Dick H. J.
err分享
err收藏
The rod sensitivity of dark adapted human infants
err2009-07-02
err0
PREAI
errAnne B. Fulton; Ronald M. Hansen
err分享
err收藏
学者 查看更多内容