arrow
Return

G-Hadoop: MapReduce across distributed data centers for data-intensive computing

delete2013-03-01
delete261
PRE
AI
L
Lizhe Wang *
J
Jie Tao
R
Rajiv Ranjan
H
H. Marten
A
Achim Streit
J
Jingying Chen
陈丹 (Dan Chen)
DOI:10.1016/j.future.2012.09.001delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recently, the computational requirements for large-scale data-intensive analysis of scientific data have grown significantly. In High Energy Physics (HEP) for example, the Large Hadron Collider (LHC) produced 13 petabytes of data in 2010. This huge amount of data is processed on more than 140 computing centers distributed across 34 countries. The MapReduce paradigm has emerged as a highly successful programming model for large-scale data-intensive computing applications. However, current MapReduce implementations are developed to operate on single cluster environments and cannot be leveraged for large-scale distributed data processing across multiple clusters. On the other hand, workflow systems are used for distributed data processing across data centers. It has been reported that the workflow paradigm has some limitations for distributed data processing, such as reliability and efficiency. In this paper, we present the design and implementation of G-Hadoop, a MapReduce framework that aims to enable large-scale distributed computing across multiple clusters. (C) 2012 Elsevier B.V. All rights reserved.
Keywords:
Cloud computing
Massive data processing
Data-intensive computing
Hadoop
MapReduce

Journal

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
Papers:
6.8K
Citations:
2.3W

Organization

K
karlsruhe institute of technology
Scholars:
2.0W
Papers: 1.4W
Citations: 23
C
China University of Geosciences
Scholars:
3.7W
Papers: 2.8W
Citations: 4.3W
H
Helmholtz Association
Scholars:
13.2W
Papers: 10.7W
Citations: 145
researcher View more organizations