arrow
返回

Efficient median based clustering and classification techniques for protein sequences

delete2006-08-22
delete7
PRE
AI
P
P. Vijaya *
M
M. Narasimha Murty
D
D.K. Subramanian
DOI:10.1007/s10044-006-0040-zdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In this paper, an efficient K-medians clustering (unsupervised) algorithm for prototype selection and Supervised K-medians (SKM) classification technique for protein sequences are presented. For sequence data sets, a median string/sequence can be used as the cluster/group representative. In K-medians clustering technique, a desired number of clusters, K, each represented by a median string/sequence, is generated and these median sequences are used as prototypes for classifying the new/test sequence whereas in SKM classification technique, median sequence in each group/class of labelled protein sequences is determined and the set of median sequences is used as prototypes for classification purpose. It is found that the K-medians clustering technique outperforms the leader based technique and also SKM classification technique performs better than that of motifs based approach for the data sets used. We further use a simple technique to reduce time and space requirements during protein sequence clustering and classification. During training and testing phase, the similarity score value between a pair of sequences is determined by selecting a portion of the sequence instead of the entire sequence. It is like selecting a subset of features for sequence data sets. The experimental results of the proposed method on K-medians, SKM and Nearest Neighbour Classifier (NNC) techniques show that the Classification Accuracy (CA) using the prototypes generated/used does not degrade much but the training and testing time are reduced significantly. Thus the experimental results indicate that the similarity score does not need to be calculated by considering the entire length of the sequence for achieving a good CA. Even space requirement is reduced during both training and classification.
Keyword:
clustering
protein sequences
median strings
sequences
set median
prototypes
feature selection
classification accuracy

期刊

Pattern Analysis and Applications 封面图
Pattern Analysis and Applications
IF:
2
论文数:
1.9K
被引数:
1.9K

机构

暂无机构信息
引用论文

引用论文

Telefon-Triage und klinische Ersteinschätzung in der Notfallmedizin zur Patientensteuerung
err2019-08-08
err0
PREAI
errB. Kumle; A. Hirschfeld-Warneken; I. Darnhofer; H. J. Busch
err分享
err收藏
err分享
err收藏
Dynamics of plant organic matter decomposition in different agricultural landscapes
err2023-03-01
err0
errOAAI
errJoão H. C. S. Silva; Alex da S. Barbosa; Daniel da S. Gomes; Italo de S. Aquino; Janaína R. da Silva
err分享
err收藏
The Development of Sexual Aggression through the Life Span
err2006-01-24
err0
PREAI
errHOWARD E. BARBAREE; RAY BLANCHARD; CALVIN M. LANGTON
err分享
err收藏
Clustering data streams: Theory and practice
err2003-05-01
err516
PREAI
errGuha, S; Meyerson, A; Mishra, N; Motwani, R; O'Callaghan, L
err分享
err收藏
Remittances and financial development: empirical evidence from heterogeneous panel of countries
err2018-02-24
err0
errOAAI
errMita Bhattacharya; John Inekwe; Sudharshan Reddy Paramati
err分享
err收藏
Vector Bundles on Complex Projective Spaces
err
IF0
err1980-01-01
err0
PREAI
errChristian Okonek; Michael Schneider; Heinz Spindler
err分享
err收藏
学者 查看更多内容