arrow
返回

Boosting phosphorylation site prediction with sequence feature-based machine learning

delete2019-08-22
delete13
PRE
AI
S
Shyantani Maiti
A
Atif Hassan
P
Pralay Mitra *
DOI:10.1002/prot.25801delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Protein phosphorylation is one of the essential posttranslation modifications playing a vital role in the regulation of many fundamental cellular processes. We propose a LightGBM-based computational approach that uses evolutionary, geometric, sequence environment, and amino acid-specific features to decipher phosphate binding sites from a protein sequence. Our method, while compared with other existing methods on 2429 protein sequences taken from standard Phospho.ELM (P.ELM) benchmark data set featuring 11 organisms reports a higher F-1 score = 0.504 (harmonic mean of the precision and recall) and ROC AUC = 0.836 (area under the curve of the receiver operating characteristics). The computation time of our proposed approach is much less than that of the recently developed deep learning-based framework. Structural analysis on selected protein sequences informs that our prediction is the superset of the phosphorylation sites, as mentioned in P.ELM data set. The foundation of our scheme is manual feature engineering and a decision tree-based classification. Hence, it is intuitive, and one can interpret the final tree as a set of rules resulting in a deeper understanding of the relationships between biophysical features and phosphorylation sites. Our innovative problem transformation method permits more control over precision and recall as is demonstrated by the fact that if we incorporate output probability of the existing deep learning framework as an additional feature, then our prediction improves (F-1 score = 0.546; ROC AUC = 0.849). The implementation of our method can be accessed at and is mirrored at .
Keyword:
computational prediction
decision tree
evolutionary information
phosphorylation sites
protein sequence analysis
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

P
Proteins Structure Function and Bioinformatics
IF:
2.8
论文数:
6.6K
被引数:
1.4W

机构

I
indian institute of technology system (iit system)
学者数:
9.5W
论文数: 9.9W
被引数: 93
引用论文

引用论文

The I-TASSER Suite: protein structure and function prediction
err2014-12-30
err4.9K
errOAAI
errYang, Jianyi; Yan, Renxiang; Roy, Ambrish; Xu, Dong; Poisson, Jonathan; Zhang, Yang
err分享
err收藏
err分享
err收藏
err分享
err收藏
Improving the accuracy of PSI-BLAST protein database searches with composition-based statistics and other refinements
err2001-07-15
err1.3K
errOAAI
errSchäffer, AA; Aravind, L; Madden, TL; Shavirin, S; Spouge, JL; Wolf, YI; Koonin, EV; Altschul, SF
err分享
err收藏
err分享
err收藏
Focused natural product elucidation by prioritizing high-throughput metabolomic studies with machine learning
err
IF0
err2019-01-31
err0
errOAAI
errNicholas J. Tobias; César Parra-Rojas; Yan-Ni Shi; Yi-Ming Shi; Svenja Simonyi; Aunchalee Thanwisai; Apichat Vitta; Narisara Chantratita; Esteban A. Hernandez-Vargas; Helge B. Bode
err分享
err收藏
Musite, a Tool for Global Prediction of General and Kinase-specific Phosphorylation Sites
err2010-12-01
err245
errOAAI
errGao, Jianjiong; Thelen, Jay J.; Dunker, A. Keith; Xu, Dong
err分享
err收藏
学者 查看更多内容