arrow
Return

Data Cube Materialization and Mining over MapReduce

delete2012-10-01
delete29
delete
OA
AI
N
Nandi, Arnab *
C
Cong Yu
R
Raghu Ramakrishnan
DOI:10.1109/TKDE.2011.257delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Computing interesting measures for data cubes and subsequent mining of interesting cube groups over massive data sets are critical for many important analyses done in the real world. Previous studies have focused on algebraic measures such as SUM that are amenable to parallel computation and can easily benefit from the recent advancement of parallel computing infrastructure such as MapReduce. Dealing with holistic measures such as TOP-K, however, is nontrivial. In this paper, we detail real-world challenges in cube materialization and mining tasks on web-scale data sets. Specifically, we identify an important subset of holistic measures and introduce MR-Cube, a MapReduce-based framework for efficient cube computation and identification of interesting cube groups on holistic measures. We provide extensive experimental analyses over both real and synthetic data. We demonstrate that, unlike existing techniques which cannot scale to the 100 million tuple mark for our data sets, MR-Cube successfully and efficiently computes cubes with holistic measures over billion-tuple data sets.
Keywords:
Data cube
cube materialization
cube mining
MapReduce
holistic measures

Journal

IEEE Transactions on Knowledge and Data Engineering cover
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
Papers:
6.8K
Citations:
3.2W

Organization

U
University System of Ohio
Scholars:
15.4W
Papers: 13.0W
Citations: 200
F
facebook inc
Scholars:
588
Papers: 381
Citations: 0
O
Ohio State University
Scholars:
4.1W
Papers: 3.2W
Citations: 80
G
Google Incorporated
Scholars:
3.5K
Papers: 1.8K
Citations: 8
researcher View more organizations