arrow
Return

Balancing reducer workload for skewed data using sampling-based partitioning

delete2014-02-01
delete12
PRE
AI
Y
Yujie Xu
W
Wenyu Qu *
Z
Zhiyang Li
Z
Zhaobin Liu
李元元 cover
李元元 (Yuanyuan Li)
H
Haifeng Li
DOI:10.1016/j.compeleceng.2013.07.001delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
MapReduce has emerged as a popular tool for distributed processing of massive data. However, it is not efficient when handling skewed data and it often leads to reducer load imbalance. In this paper, we address the problem of how to efficiently partition intermediate keys to balance the workload of all reducers when processing skewed data. We present a sampling scheme to compute the approximate distribution of key frequency, estimate the overall distribution and then make a partition scheme in advance. Then, we apply it to map phase of the executing MapReduce job. This work not only provides a load-balanced partition strategy, but also keeps a high performance of synchronous mode of MapReduce. We also propose two partition methods based on sampling results: cluster combination and cluster split combination. The experimental results show that our methods achieve a better time and load balancing results. (C) 2013 Elsevier Ltd. All rights reserved.

Journal

C
Computers and Electrical Engineering
IF:
4.9
Papers:
6.7K
Citations:
1.3W

Organization

D
Dalian Maritime University
Scholars:
1.2W
Papers: 7.8K
Citations: 6.3K