arrow
Return

A fast and resource efficient mining algorithm for discovering frequent patterns in distributed computing environments

delete2015-11-01
delete16
PRE
AI
K
Kawuu W. Lin *
S
Sheng-Hao Chung
DOI:10.1016/j.future.2015.05.009delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The advancement of electronic technology enables us to collect logs from various devices. Such logs require detailed analysis in order to be broadly useful. Data mining is a technique that has been widely used to extract hidden information from such data. Data mining is mainly composed of association rules mining, sequent pattern mining, classification and clustering. Association rules mining has attracted significant attention and been successfully applied to various fields. Although the past studies can effectively discover frequent patterns to deduce association rules, execution efficiency is still a critical problem. To speed up execution, many methods using parallel and distributed computing technology have been proposed in recent years. Most of the past studies focused on parallelizing the workload in a high end machine or in distributed computing environments like grid or cloud computing systems; however, very few of them discuss how to efficiently determine the appropriate number of computing nodes, considering execution efficiency and load balancing. An intuition is that execution speed is proportional to the number of computing nodes that is, more the number of computing nodes, faster is the execution speed. However, this is incorrect for such algorithms because of the inherently algorithmic design. Allocating too many computing nodes can lead to high execution time. In addition to the execution inefficiency, inappropriate resource allocation is a waste of computing power and network bandwidth. At the same time, load cannot be effectively distributed if there are too few nodes allocated. In this paper, we propose a fast, load balancing and resource efficient algorithm named FLR-Mining for discovering frequent patterns in distributed computing systems. FLR-Mining is capable of determining the appropriate number of computing nodes automatically and achieving better load balancing as compared with existing methods. Through empirical evaluation, FLR-Mining is shown to deliver excellent performance in terms of execution efficiency and load balancing. (C) 2015 Elsevier B.V. All rights reserved.
Keywords:
Data mining
Frequent pattern mining
Distributed mining
Parallel mining
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
Papers:
6.8K
Citations:
2.3W

Organization

N
national kaohsiung university of science & technology
Scholars:
4.3K
Papers: 4.8K
Citations: 3
Cited Papers

Cited Papers

Adjuvant Immunotherapy with B.C.G. in Carcinoma of the Prostate
err1977-06-01
err0
PREAI
errM. R. G. ROBINSON; C. C. RIGBY; R. C. B. PUGH; D. C. DUMONDE
errShare
errSave
Spatial Genetic Structure of Campanula sabatia, a Threatened Narrow Endemic Species of the Mediterranean Basin
err2012-07-21
err0
PREAI
errFederica Nicoletti; Laura De Benedetti; Marcello Airò; Barbara Ruffoni; Antonio Mercuri; Luigi Minuto; Gabriele Casazza
errShare
errSave
Data Mining with Big Data
err2014-01-01
err1.9K
PREAI
errWu, Xindong; Zhu, Xingquan; Wu, Gong-Qing; Ding, Wei
errShare
errSave
Gene association analysis: a survey of frequent pattern mining from gene expression data
err2009-10-08
err58
errOAAI
errAlves, Ronnie; Rodriguez-Baena, Domingo S.; Aguilar-Ruiz, Jesus S.
errShare
errSave
researcher View more