返回
Big SQL systems: an experimental evaluation
DOI:10.1007/s10586-019-02914-4.png)
摘要
En 中文
Recently, Big Data systems have been gaining increasing popularity on handling the massive amounts of data that are continuously generated in our digital world. While the Hadoop framework has pioneered the area of Big Data processing systems, it had clear performance limitations on providing the best performance of processing massive amounts of structured data. In addition, practically, many users of the big data systems face some challenges on dealing with the APIs and the low level programming abstractions of the Big Data System and they would prefer to use SQL (in which they are more proficient) as a high-level declarative language to express their tasks while leaving all of the execution optimization details to the backend engine. Thus, several systems have been designed and implemented to tackle these challenges by designing and implementing scalable query execution engines for processing massive structured data while supporting SQL interfaces. In this article, we present an extensive experimental study of four popular systems in this domain, namely, Apache Hive, SPARK SQL, Apache Impala and PrestoDB. In particular, we report and analyze the performance characteristics of these systems using three different benchmarks, namely, TPC-H, TPC-DS and TPCx-BB. Finally, we report a set of insights and important lessons that we have learned from conducting our experiments.
Keyword:
Big data
Big SQL
Benchmarking
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
C
IF:
4.1
论文数:
5.0K
被引数:
7.5K
机构
暂无机构信息
引用论文
Primary structure of the two (4 Fe-4 S) clusters ferredoxin from Desulfovibrio desulfuricans (strain Norway 4)
Biochimie
IF0
Capacity optimization of hybrid renewable energy system considering part-load ratio and resource endowment考虑部分负荷率和资源禀赋的混合可再生能源系统容量优化

