arrow
Return

Parallel Clustering Algorithm for Large Data Sets with Applications in Bioinformatics

delete2009-04-01
delete50
PRE
AI
V
Victor Olman *
F
Fenglou Mao
H
Hongwei Wu
徐鹰 cover
徐鹰 (Ying Xu)
DOI:10.1109/TCBB.2007.70272delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large sets of bioinformatical data provide a challenge in time consumption while solving the cluster identification problem, and that is why a parallel algorithm is so needed for identifying dense clusters in a noisy background. Our algorithm works on a graph representation of the data set to be analyzed. It identifies clusters through the identification of densely intraconnected subgraphs. We have employed a minimum spanning tree (MST) representation of the graph and solve the cluster identification problem using this representation. The computational bottleneck of our algorithm is the construction of an MST of a graph, for which a parallel algorithm is employed. Our high-level strategy for the parallel MST construction algorithm is to first partition the graph, then construct MSTs for the partitioned subgraphs and auxiliary bipartite graphs based on the subgraphs, and finally merge these MSTs to derive an MST of the original graph. The computational results indicate that when running on 150 CPUs, our algorithm can solve a cluster identification problem on a data set with 1,000,000 data points almost 100 times faster than on single CPU, indicating that this program is capable of handling very large data clustering problems in an efficient manner. We have implemented the clustering algorithm as the software CLUMP.
Keywords:
Pattern recognition
clustering algorithm
genome application
parallel processing

Journal

I
IEEE-ACM Transactions on Computational Biology and Bioinformatics
IF:
3.4
Papers:
3.3K
Citations:
6.4K

Organization

U
university system of georgia
Scholars:
7.3W
Papers: 6.6W
Citations: 101
Cited Papers

Cited Papers

An efficient algorithm for large-scale detection of protein families
err2002-04-01
err2.9K
errOAAI
errEnright, AJ; Van Dongen, S; Ouzounis, CA
errShare
errSave
errShare
errSave
The COG database: new developments in phylogenetic classification of proteins from complete genomes
err2001-01-01
err1.7K
errOAAI
errTatusov, RL; Natale, DA; Garkavtsev, IV; Tatusova, TA; Shankavaram, UT; Rao, BS; Kiryutin, B; Galperin, MY; Fedorova, ND; Koonin, EV
errShare
errSave
IP over ICN - The better IP?
err2015-06-01
err0
PREAI
errDirk Trossen; Martin J. Reed; Janne Riihijarvi; Michael Georgiades; Nikos Fotiou; George Xylomenos
errShare
errSave
Hierarchical classification of functionally equivalent genes in prokaryotes
err2007-03-11
err10
errOAAI
errWu, Hongwei; Mao, Fenglou; Olman, Victor; Xu, Ying
errShare
errSave
researcher View more