返回
An iterative sampling method for online aggregation
DOI:10.1007/s10586-017-1451-x.png)
摘要
En 中文
Online aggregation (OLA) makes it possible to save cost by taking acceptable approximate early answers. Compared to the precise results, computing the approximate ones are more cost effective, especially for large-scale datasets. The user can terminate the processing at any time, when he/she is satisfied with the quality of the result. And the performance of OLA relies on the sampling approach and estimation model. But in large scale distributed computing environment, how to realize OLA more efficiently is a challenging problem. In this paper, we consider the problem of providing OLA in the distributed computing environment and propose a Hadoop-based iterative sampling method for online aggregation. The desired precision of the user can be met by two iteration samplings. To avoid the effects of data bias, we propose a layered sampling method to ensure that the approximate aggregation result is statistically meaningful. The experimental results showed the layered sampling method considers not only the time efficiency, but also the usage of computing and storage resources of Hadoop.
Keyword:
Online aggregation
Iteration
Sampling
Query processing
Hadoop
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
C
IF:
4.1
论文数:
5.1K
被引数:
7.5K
机构
引用论文
An intriguing N-oxide-functionalized 3D flexible microporous MOF exhibiting highly selectivity for CO2 with a gate effect
Polyhedron
IF0
没有更多内容

