arrow
Return

Big data processing tools: An experimental performance evaluation

delete2018-12-25
delete13
PRE
AI
M
M.R.D. Rodrigues
M
Maribel Yasmina Santos
J
Jorge Bernardino *
DOI:10.1002/widm.1297delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Big Data is currently a hot topic of research and development across several business areas mainly due to recent innovations in information and communication technologies. One of the main challenges of Big Data relates to how one should efficiently handle massive volumes of complex data. Due to the notorious complexity of the data that can be collected from multiple sources, usually motivated by increasing data volumes gathered at high velocity, efficient processing mechanisms are needed for data analysis purposes. Motivated by the rapid growth in technology, development of tools, and frameworks for Big Data, there is much discussion about Big Data querying tools and, specifically, those that are more appropriated for specific analytical needs. This paper describes and evaluates the following popular Big Data processing tools: Drill, HAWQ, Hive, Impala, Presto, and Spark. An experimental evaluation using the Transaction Processing Council (TPC-H) benchmark is presented and discussed, highlighting the performance of each tool, according to different workloads and query types.
Keywords:
Big Data
Big Data analytics
query processing
SQL-on-Hadoop
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Wiley Interdisciplinary Reviews-Data Mining and Knowledge Discovery cover
Wiley Interdisciplinary Reviews-Data Mining and Knowledge Discovery
IF:
11.7
Papers:
534
Citations:
5.3K

Organization

I
instituto politecnico de coimbra (ipc)
Scholars:
420
Papers: 362
Citations: 1
U
universidade do minho
Scholars:
1.1W
Papers: 1.1W
Citations: 10