返回
Automatic Classification of Protein Structures Using Physicochemical Parameters
DOI:10.1007/s12539-013-0199-0.png)
摘要
En 中文
Protein classification is the first step to functional annotation; SCOP and Pfam databases are currently the most relevant protein classification schemes. However, the disproportion in the number of three dimensional (3D) protein structures generated versus their classification into relevant superfamilies/families emphasizes the need for automated classification schemes. Predicting function of novel proteins based on sequence information alone has proven to be a major challenge. The present study focuses on the use of physicochemical parameters in conjunction with machine learning algorithms (Naive Bayes, Decision Trees, Random Forest and Support Vector Machines) to classify proteins into their respective SCOP superfamily/Pfam family, using sequence derived information. Spectrophores (TM), a 1D descriptor of the 3D molecular field surrounding a structure was used as a benchmark to compare the performance of the physicochemical parameters. The machine learning algorithms were modified to select features based on information gain for each SCOP superfamily/Pfam family. The effect of combining physicochemical parameters and spectrophores on classification accuracy (CA) was studied. Machine learning algorithms trained with the physicochemical parameters consistently classified SCOP superfamilies and Pfam families with a classification accuracy above 90%, while spectrophores performed with a CA of around 85%. Feature selection improved classification accuracy for both physicochemical parameters and spectrophores based machine learning algorithms. Combining both attributes resulted in a marginal loss of performance. Physicochemical parameters were able to classify proteins from both schemes with classification accuracy ranging from 90-96%. These results suggest the usefulness of this method in classifying proteins from amino acid sequences.
Keyword:
protein classification
machine learning algorithms
physicochemical parameters
feature selection
svm
random forest
naive bayes
decision tree
SCOP classification
Pfam classification
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
I
IF:
3.9
论文数:
955
被引数:
1.5K
机构
引用论文
The HHpred interactive server for protein homology detection and structure prediction用于蛋白质同源性检测和结构预测的HHpred交互式服务器
NUCLEIC ACIDS RESEARCH
IF13.1
Peperomia at SemEval-2018 Task 2: Vector Similarity Based Approach
for Emoji PredictionPeperomia 在 SemEval-2018 任务 2:基于向量相似度的表情符号预测方法
EMBOSS: The European molecular biology open software suiteEMBOSS: 欧洲分子生物学开放软件套件
TRENDS IN GENETICS
IF16.3

