Return
Answering Subset Query Over Multi-Attribute Data Streams Using Hyper-USS
DOI:10.1109/TKDE.2025.3611502.png)
Abstract
En 中文
Approximate queries offer an efficient means of analyzing massive data streams under acceptable errors. Among these, subset queries over multiple attributes are common in many real-world applications. While sketches offer promising approximate solutions for massive data streams, efficiently supporting subset queries over multiple statistical attributes remains a significant challenge. To address this, we propose Hyper-USS, a novel sketching solution that accurately and efficiently supports subset queries over data streams involving multiple statistical attributes. With Joint Variance Optimization, Hyper-USS provides unbiased estimation and optimizes estimation variance jointly, addressing the challenge of accurately estimating multiple statistical attributes in the sketch design. The algorithm records the information of keys and all attributes in one sketch, ensuring high insertion efficiency. Furthermore, its three speed-optimized versions are introduced to handle the growing number of statistical attributes in data streams. Experimental results show that Hyper-USS and its three speed-optimized versions consistently surpass state-of-the-art methods that support subset queries in both estimation accuracy and insertion throughput. Specifically, Hyper-USS improves accuracy by at least 38%, while the algorithm and its three speed-optimized versions achieve throughput improvements of up to <inline-formula><tex-math notation="LaTeX">$31.90\times$</tex-math></inline-formula>, <inline-formula><tex-math notation="LaTeX">$45.31\times$</tex-math></inline-formula>, <inline-formula><tex-math notation="LaTeX">$49.21\times$</tex-math></inline-formula>, and <inline-formula><tex-math notation="LaTeX">$58.03\times$</tex-math></inline-formula>, respectively.
Keywords:
Sketch
multi-attribute data streams
subset query
unbiased estimation
Journal
IF:
10.4
Papers:
6.7K
Citations:
3.2W

