arrow
返回

Parallel Star Join plus DataIndexes: Efficient query processing in data warehouses and OLAP

delete2002-11-01
delete13
PRE
AI
A
Anindya Datta *
D
Debra VanderMeer
K
Krithi Ramamritham
DOI:10.1109/TKDE.2002.1047769delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
On-Line Analytical Processing (OLAP) refers to the technologies that allow users to efficiently retrieve data from the data warehouse for decision-support purposes. Data warehouses tend to be extremely large-it is quite possible for a data warehouse to be hundreds of gigabytes to terabytes in size [3]. Queries tend to be complex and ad hoc, often requiring computationally expensive operations such as joins and aggregation. Given this, we are interested in developing strategies for improving query processing in data warehouses by exploring the applicability of parallel processing techniques. In particular, we exploit the natural partitionability of a star schema and render it even more efficient by applying DataIndexes-a storage structure that serves both as an index as well as data and lends itself naturally to vertical partitioning of the data. Dataindexes are derived from the various special purpose access mechanisms currently supported in commercial OLAP products. Specifically, we propose a declustering strategy which incorporates both task and data partitioning and present the Parallel Star Join (PSJ) Algorithm, which provides a means to perform a star join in parallel using efficient operations involving only rowsets and projection columns. We compare the performance of the PSJ Algorithm with two parallel query processing strategies. The first is a parallel join strategy utilizing the Bitmap Join Index (BJI), arguably the state-of-the-art OLAP join structure in use today. For the second strategy we choose a well-known parallel join algorithm, namely the pipelined hash algorithm. To assist in the performance comparison, we first develop a cost model of the disk access and transmission costs for all three approaches. Performance comparisons show that the Dataindex-based approach leads to dramatically lower disk access costs than the BJI, as well as the hybrid hash approaches, in both speedup and scaleup experiments, while the hash-based approach outperforms the BJI in disk access costs. With regard to transmission overhead, our performance results show that PSJ and BJI outperform the hash-based approach. Overall, our parallel star join algorithm and dataindexes form a winning combination.
Keyword:
parallel star join
OLAP
query processing
dataindexes
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Knowledge and Data Engineering 封面图
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
论文数:
6.8K
被引数:
3.2W

机构

暂无机构信息
引用论文

引用论文

Preventative effect of repeated nasal applications of capsaicin in cluster headache
errPain
IF0
err1994-12-01
err0
PREAI
errBruno M. Fusco; Simone Marabini; Carlo A. Maggi; Giuseppe Fiore; Pierangelo Geppetti
err分享
err收藏
Single agent activity of U3-1402, a HER3-targeting antibody-drug conjugate, in HER3-overexpressing metastatic breast cancer: Updated results from a phase I/II trial
err2019-05-01
err0
errOAAI
errK. Yonemori; N. Masuda; S. Takahashi; T. Kogawa; T. Nakayama; Y. Yamamoto; M. Takahashi; T. Toyama; T. Saeki; H. Iwata
err分享
err收藏
Meibomian Gland Dysfunction Associated With Periocular Radiotherapy
err2017-09-08
err0
PREAI
errYoung Jun Woo; JaeSang Ko; Yong Woo Ji; Tae-im Kim; Jin Sook Yoon
err分享
err收藏
学者 查看更多内容