arrow
Return

An iterative sampling method for online aggregation

delete2017-12-19
delete0
PRE
AI
张志强 (Zhiqiang Zhang) *
J
Jianghua Hu
X
Xiaoqin Xie
H
Haiwei Pan
X
Xiaoning Feng
DOI:10.1007/s10586-017-1451-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Online aggregation (OLA) makes it possible to save cost by taking acceptable approximate early answers. Compared to the precise results, computing the approximate ones are more cost effective, especially for large-scale datasets. The user can terminate the processing at any time, when he/she is satisfied with the quality of the result. And the performance of OLA relies on the sampling approach and estimation model. But in large scale distributed computing environment, how to realize OLA more efficiently is a challenging problem. In this paper, we consider the problem of providing OLA in the distributed computing environment and propose a Hadoop-based iterative sampling method for online aggregation. The desired precision of the user can be met by two iteration samplings. To avoid the effects of data bias, we propose a layered sampling method to ensure that the approximate aggregation result is statistically meaningful. The experimental results showed the layered sampling method considers not only the time efficiency, but also the usage of computing and storage resources of Hadoop.
Keywords:
Online aggregation
Iteration
Sampling
Query processing
Hadoop
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

C
Cluster Computing-The Journal of Networks Software Tools and Applications
IF:
4.1
Papers:
5.0K
Citations:
7.5K

Organization

H
Harbin Engineering University
Scholars:
1.9W
Papers: 1.3W
Citations: 1.3W