arrow
Return

End-to-End Optimization for Geo-Distributed MapReduce

delete2016-07-01
delete38
delete
OA
AI
B
Benjamin Heintz *
A
Abhishek Chandra
R
Ramesh K. Sitaraman
J
Jon Weissman
DOI:10.1109/TCC.2014.2355225delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
MapReduce has proven remarkably effective for a wide variety of data-intensive applications, but it was designed to run on large single-site homogeneous clusters. Researchers have begun to explore the extent to which the original MapReduce assumptions can be relaxed, including skewed workloads, iterative applications, and heterogeneous computing environments. This paper continues this exploration by applying MapReduce across geo-distributed data over geo-distributed computation resources. Using Hadoop, we show that network and node heterogeneity and the lack of data locality lead to poor performance, because the interaction of MapReduce phases becomes pronounced in the presence of heterogeneous network behavior. To address these problems, we take a two-pronged approach: We first develop a model-driven optimization that serves as an oracle, providing high-level insights. We then apply these insights to design cross-phase optimization techniques that we implement and demonstrate in a real-world MapReduce implementation. Experimental results in both Amazon EC2 and PlanetLab show the potential of these techniques as performance is improved by 7-18 percent depending on the execution environment and application.
Keywords:
Batch processing systems
distributed systems
parallel systems
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

I
IEEE Transactions on Cloud Computing
IF:
5
Papers:
1.8K
Citations:
4.3K

Organization

U
University of Minnesota Twin Cities
Scholars:
3.7W
Papers: 3.1W
Citations: 58