arrow
返回

Frequent Itemsets Mining for Big Data: A Comparative Analysis

delete2017-09-01
delete32
delete
OA
AI
D
Daniele Apiletti
E
Elena Baralis
T
Tania Cerquitelli
P
Paolo Garza
F
Fabio Pulvirenti *
L
Luca Venturini
DOI:10.1016/j.bdr.2017.06.006delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Itemset mining is a well-known exploratory data mining technique used to discover interesting correlations hidden in a data collection. Since it supports different targeted analyses, it is profitably exploited in a wide range of different domains, ranging from network traffic data to medical records. With the increasing amount of generated data, different scalable algorithms have been developed, exploiting the advantages of distributed computing frameworks, such as Apache Hadoop and Spark. This paper reviews Hadoop-and Spark-based scalable algorithms addressing the frequent itemset mining problem in the Big Data domain through both theoretical and experimental comparative analyses. Since the itemset mining task is computationally expensive, its distribution and parallelization strategies heavily affect memory usage, load balancing, and communication costs. A detailed discussion of the algorithmic choices of the distributed methods for frequent itemset mining is followed by an experimental analysis comparing the performance of state-of-the-art distributed implementations on both synthetic and real datasets. The strengths and weaknesses of the algorithms are thoroughly discussed with respect to the dataset features (e.g., data distribution, average transaction length, number of records), and specific parameter settings. Finally, based on theoretical and experimental analyses, open research directions for the parallelization of the itemset mining problem are presented. (C) 2017 Elsevier Inc. All rights reserved.
Keyword:
Big Data
Frequent itemset mining
Hadoop and Spark platforms
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Big Data Research 封面图
Big Data Research
IF:
4.2
论文数:
406
被引数:
1.1K

机构

P
Polytechnic University of Turin
学者数:
1.3W
论文数: 1.3W
被引数: 1.3W