返回
LIBRA: Lightweight Data Skew Mitigation in MapReduce
DOI:10.1109/TPDS.2014.2350972.png)
摘要
En 中文
MapReduce is an effective tool for parallel data processing. One significant issue in practical MapReduce applications is data skew: the imbalance in the amount of data assigned to each task. This causes some tasks to take much longer to finish than others and can significantly impact performance. This paper presents LIBRA, a lightweight strategy to address the data skew problemamong the reducers of MapReduce applications. Unlike previous work, LIBRA does not require any pre-run sampling of the input data or prevent the overlap between the map and the reduce stages. It uses an innovative sampling method which can achieve a highly accurate approximation to the distribution of the intermediate data by sampling only a small fraction of the intermediate data during the normal map processing. It allows the reduce tasks to start copying as soon as the chosen sample map tasks (only a small fraction of map tasks which are issued first) complete. It supports the split of large keys when application semantics permit and the total order of the output data. It considers the heterogeneity of the computing resources when balancing the load among the reduce tasks appropriately. LIBRA is applicable to a wide range of applications and is transparent to the users. We implement LIBRA in Hadoop and our experiments show that LIBRA has negligible overhead and can speed up the execution of some popular applications by
Keyword:
MapReduce
data skew
sampling
partitioning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6
论文数:
5.2K
被引数:
1.1W
机构
引用论文
Semen Donors as the Source of Sexually Transmitted Diseases in Artificially Inseminated Women: The Saga Unfolds精液捐赠者是人工授精妇女性传播疾病的来源: 传奇故事展开
Dynamic Resource Allocation Using Virtual Machines for Cloud Computing Environment云计算环境中使用虚拟机的动态资源分配

