arrow
Return

Big Data with Cloud Computing: an insight on the computing environment, MapReduce, and programming frameworks

delete2014-09-29
delete175
PRE
AI
A
Alberto Fernández *
S
Sara del Río
V
Victoria López
M
María José del Jesús
J
José M. Benítez
F
Francisco Herrera
DOI:10.1002/widm.1134delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The term Big Data' has spread rapidly in the framework of Data Mining and Business Intelligence. This new scenario can be defined by means of those problems that cannot be effectively or efficiently addressed using the standard computing resources that we currently have. We must emphasize that Big Data does not just imply large volumes of data but also the necessity for scalability, i.e., to ensure a response in an acceptable elapsed time. When the scalability term is considered, usually traditional parallel-type solutions are contemplated, such as the Message Passing Interface or high performance and distributed Database Management Systems. Nowadays there is a new paradigm that has gained popularity over the latter due to the number of benefits it offers. This model is Cloud Computing, and among its main features we has to stress its elasticity in the use of computing resources and space, less management effort, and flexible costs. In this article, we provide an overview on the topic of Big Data, and how the current problem can be addressed from the perspective of Cloud Computing and its programming frameworks. In particular, we focus on those systems for large-scale analytics based on the MapReduce scheme and Hadoop, its open-source implementation. We identify several libraries and software projects that have been developed for aiding practitioners to address this new programming model. We also analyze the advantages and disadvantages of MapReduce, in contrast to the classical solutions in this field. Finally, we present a number of programming frameworks that have been proposed as an alternative to MapReduce, developed under the premise of solving the shortcomings of this model in certain scenarios and platforms. WIREs Data Mining Knowl Discov 2014, 4:380-409. doi: 10.1002/widm.1134 For further resources related to this article, please visit the . Conflict of interest: The authors have declared no conflicts of interest for this article.
Keywords:
DATA SCIENCE
MAP-REDUCE
PERFORMANCE
CLASSIFICATION
ALGORITHMS
CHALLENGES
TECHNOLOGIES
ASSOCIATION
SIMILARITY
DATABASES
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Wiley Interdisciplinary Reviews-Data Mining and Knowledge Discovery cover
Wiley Interdisciplinary Reviews-Data Mining and Knowledge Discovery
IF:
11.7
Papers:
537
Citations:
5.3K

Organization

K
King Abdulaziz University
Scholars:
2.0W
Papers: 1.9W
Citations: 3.3W
U
universidad de jaen
Scholars:
4.5K
Papers: 4.6K
Citations: 4
U
University of Granada
Scholars:
2.3W
Papers: 1.9W
Citations: 24
researcher View more organizations