arrow
Return

Efficient construction of histograms for multidimensional data using quad-trees

delete2011-12-01
delete2
PRE
AI
Y
Yohan Roh *
J
Jae Ho Kim
J
Jin Hyun Son
M
Myoung Ho Kim
DOI:10.1016/j.dss.2011.05.006delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Histograms can be useful in estimating the selectivity of queries in areas such as database query optimization and data exploration. In this paper, we propose a new histogram method for multidimensional data, called the Q-Histogram, based on the use of the quad-tree, which is a popular index structure for multidimensional data sets. The use of the compact representation of the target data obtainable from the quad-tree allows a fast construction of a histogram with the minimum number of scanning, i.e., only one scanning, of the underlying data. In addition to the advantage of computation time, the proposed method also provides a better performance than other existing methods with respect to the quality of selectivity estimation. We present a new measure of data skew for a histogram bucket, called the weighted bucket skew. Then, we provide an effective technique for skew-tolerant organization of histograms. Finally, we compare the accuracy and efficiency of the proposed method with other existing methods using both real-life data sets and synthetic data sets. The results of experiments show that the proposed method generally provides a better performance than other existing methods in terms of accuracy as well as computational efficiency. Crown Copyright (C) 2011 Published by Elsevier B.V. All rights reserved.
Keywords:
Data management
Query optimization
Selectivity estimation
Multidimensional histograms

Journal

Decision Support Systems cover
Decision Support Systems
IF:
6.8
Papers:
3.8K
Citations:
1.5W

Organization

S
samsung
Scholars:
8.6K
Papers: 6.4K
Citations: 8
H
hanyang university
Scholars:
2.9W
Papers: 2.7W
Citations: 36