Return
Generalized Measure-Biased Sampling and Priority Sampling
DOI:10.1109/TKDE.2023.3340673.png)
Abstract
En 中文
Query with aggregates is one of the most important classes of ad-hoc queries. Since query response time is critical in many scenarios, small errors are usually tolerable for query processing. In this work, we adopt sampling to provide fast approximate answers to distribution query and subset-sum query. On the one hand, uniform sampler is sub-optimal. On the other hand, both measure-biased sampler and priority sampler need to create a sample for each measure column. It leads to expensive storage cost, when there are dozens or hundreds of measure columns in the table. To address this issue, we generalize both measure-biased sampler and priority sampler, which can compress the samples but still provide fast approximate answers to both distribution query and subset-sum query within a user-specified error bound. Besides, we establish the relationship between measure-biased sampler and priority sampler by constructing a measure-biased sample from a priority sample. We also extend the priority sampler to support multiple types of aggregates for arbitrary subset. In the extensive experimental evaluation, our generalized samplers achieve a remarkable improvement over the original samplers in terms of the error metrics.
Keywords:
Aggregates
Measurement uncertainty
Weight measurement
Data visualization
Velocity measurement
Time factors
Time measurement
Aggregation query
measure-biased sampling
priority sampling
stratified sampling
Journal
IF:
10.4
Papers:
6.8K
Citations:
3.2W

