返回
Data mining in distributed environment: a survey
DOI:10.1002/widm.1216.png)
摘要
En 中文
Due to the rapid growth of resource sharing, distributed systems are developed, which can be used to utilize the computations. Data mining (DM) provides powerful techniques for finding meaningful and useful information from a very large amount of data, and has a wide range of real-world applications. However, traditional DM algorithms assume that the data is centrally collected, memory-resident, and static. It is challenging to manage the large-scale data and process them with very limited resources. For example, large amounts of data are quickly produced and stored at multiple locations. It becomes increasingly expensive to centralize them in a single place. Moreover, traditional DM algorithms generally have some problems and challenges, such as memory limits, low processing ability, and inadequate hard disk, and so on. To solve the above problems, DM on distributed computing environment [also called distributed data mining (DDM)] has been emerging as a valuable alternative in many applications. In this study, a survey of state-of-the-art DDM techniques is provided, including distributed frequent item-set mining, distributed frequent sequence mining, distributed frequent graph mining, distributed clustering, and privacy preserving of distributed data mining. We finally summarize the opportunities of data mining tasks in distributed environment. (C) 2017Wiley Periodicals, Inc.
Keyword:
K-MEANS
SEQUENTIAL PATTERNS
FREQUENT PATTERNS
ASSOCIATION RULES
PARALLEL
PRIVACY
EFFICIENT
ALGORITHM
ITEMSETS
GENERATION
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
11.7
论文数:
544
被引数:
5.3K

