arrow
返回

A method for multiple-sequence-alignment-free protein structure prediction using a protein language model

delete2023-10-09
delete30
delete
OA
AI
X
Xiaomin Fang
F
Fan Wang *
L
Lihang Liu
J
Jingzhou He
D
Dayong Lin
Y
Yingfei Xiang
K
Kunrui Zhu
X
Xiaonan Zhang
H
Hua Wu
李
李晖 (Hui Li)
L
Le Song *
DOI:10.1038/s42256-023-00721-6delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Protein structure prediction pipelines based on artificial intelligence, such as AlphaFold2, have achieved near-experimental accuracy. These advanced pipelines mainly rely on multiple sequence alignments (MSAs) as inputs to learn the co-evolution information from the homologous sequences. Nonetheless, searching MSAs from protein databases is time consuming, usually taking tens of minutes. Consequently, we attempt to explore the limits of fast protein structure prediction by using only primary structures of proteins. Our proposed method, HelixFold-Single, combines a large-scale protein language model with the superior geometric learning capability of AlphaFold2. HelixFold-Single first pre-trains a large-scale protein language model with thousands of millions of primary structures utilizing the self-supervised learning paradigm, which will be used as an alternative to MSAs for learning the co-evolution information. Then, by combining the pre-trained protein language model and the essential components of AlphaFold2, we obtain an end-to-end differentiable model to predict the three-dimensional coordinates of atoms from only the primary structure. HelixFold-Single is validated on datasets CASP14 and CAMEO, achieving competitive accuracy with the MSA-based methods on targets with large homologous families. Furthermore, HelixFold-Single consumes much less time than the mainstream pipelines for protein structure prediction, demonstrating its potential in tasks requiring many predictions. AlphaFold2 has revolutionized bioinformatics, but its ability to predict protein structures with high accuracy comes at the price of a costly database search for multiple sequence alignments. Fang and colleagues pre-train a large-scale protein language model and use it in conjunction with AlphaFold2 as a fully trainable and efficient model for structure prediction.

期刊

Nature Machine Intelligence 封面图
Nature Machine Intelligence
IF:
23.9
论文数:
1.3K
被引数:
1.5W

机构

B
baidu
学者数:
578
论文数: 471
被引数: 1
引用论文

引用论文

Opioid activities of β-casomorphinsΒ-酪啡肽的阿片活性
err1981-04-01
err0
PREAI
errVictor Brantl; Hansjörg Teschemacher; Julia Bläsig; Agnes Henschen; Friedrich Lottspeich
err分享
err收藏
The I-TASSER Suite: protein structure and function prediction
err2014-12-30
err4.9K
errOAAI
errYang, Jianyi; Yan, Renxiang; Roy, Ambrish; Xu, Dong; Poisson, Jonathan; Zhang, Yang
err分享
err收藏
Continuous Automated Model EvaluatiOn (CAMEO)-Perspectives on the future of fully automated evaluation of structure prediction methods
err2021-08-19
err35
errOAAI
errRobin, Xavier; Haas, Juergen; Gumienny, Rafal; Smolinski, Anna; Tauriello, Gerardo; Schwede, Torsten
err分享
err收藏
RCSB Protein Data Bank: powerful new tools for exploring 3D structures of biological macromolecules for basic and applied research and education in fundamental biology, biomedicine, biotechnology, bioengineering and energy sciencesRCSB蛋白质数据库: 强大的新工具,用于探索生物大分子的3D结构,用于基础生物学,生物医学,生物技术,生物工程和能源科学的基础和应用研究和教育
err2020-11-19
err994
errOAAI
errBurley, Stephen K.; Bhikadiya, Charmi; Bi, Chunxiao; Bittrich, Sebastian; Chen, Li; Crichlow, Gregg, V; Christie, Cole H.; Dalenberg, Kenneth; Di Costanzo, Luigi; Duarte, Jose M.; Dutta, Shuchismita; Feng, Zukang; Ganesan, Sai; Goodsell, David S.; Ghosh, Sutapa; Green, Rachel Kramer; Guranovic, Vladimir; Guzenko, Dmytro; Hudson, Brian P.; Lawson, Catherine L.; Liang, Yuhe; Lowe, Robert; Namkoong, Harry; Peisach, Ezra; Persikova, Irina; Randle, Chris; Rose, Alexander; Rose, Yana; Sali, Andrej; Segura, Joan; Sekharan, Monica; Shao, Chenghua; Tao, Yi-Ping; Voigt, Maria; Westbrook, John D.; Young, Jasmine Y.; Zardecki, Christine; Zhuravleva, Marina
err分享
err收藏
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences生物结构和功能源于扩展无监督学习以2.5亿蛋白质序列
err2021-04-05
err1.1K
errOAAI
errRives, Alexander; Meier, Joshua; Sercu, Tom; Goyal, Siddharth; Lin, Zeming; Liu, Jason; Guo, Demi; Ott, Myle; Zitnick, C. Lawrence; Ma, Jerry; Fergus, Rob
err分享
err收藏
The Protein Data Bank蛋白质数据库
err2000-01-01
err3.3W
errOAAI
errBerman, HM; Westbrook, J; Feng, Z; Gilliland, G; Bhat, TN; Weissig, H; Shindyalov, IN; Bourne, PE
err分享
err收藏
学者 查看更多内容