arrow
Return

Generalized Measure-Biased Sampling and Priority Sampling

delete2024-11-01
delete0
PRE
AI
Z
Zhao Chang
F
Feifei Li
Y
Yulong Shen *
DOI:10.1109/TKDE.2023.3340673delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Query with aggregates is one of the most important classes of ad-hoc queries. Since query response time is critical in many scenarios, small errors are usually tolerable for query processing. In this work, we adopt sampling to provide fast approximate answers to distribution query and subset-sum query. On the one hand, uniform sampler is sub-optimal. On the other hand, both measure-biased sampler and priority sampler need to create a sample for each measure column. It leads to expensive storage cost, when there are dozens or hundreds of measure columns in the table. To address this issue, we generalize both measure-biased sampler and priority sampler, which can compress the samples but still provide fast approximate answers to both distribution query and subset-sum query within a user-specified error bound. Besides, we establish the relationship between measure-biased sampler and priority sampler by constructing a measure-biased sample from a priority sample. We also extend the priority sampler to support multiple types of aggregates for arbitrary subset. In the extensive experimental evaluation, our generalized samplers achieve a remarkable improvement over the original samplers in terms of the error metrics.
Keywords:
Aggregates
Measurement uncertainty
Weight measurement
Data visualization
Velocity measurement
Time factors
Time measurement
Aggregation query
measure-biased sampling
priority sampling
stratified sampling

Journal

IEEE Transactions on Knowledge and Data Engineering cover
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
Papers:
6.8K
Citations:
3.2W

Organization

A
alibaba group
Scholars:
1.1K
Papers: 789
Citations: 0
X
Xidian University
Scholars:
2.4W
Papers: 1.9W
Citations: 9.7K