arrow
返回

BigBind: Learning from Nonstructural Data for Structure-Based Virtual Screening

delete2023-12-19
delete3
PRE
AI
M
Michael Brocidiacono *
P
Paul Francoeur
R
Rishal Aggarwal
K
Konstantin I. Popov
D
David Ryan Koes
A
Alexander Tropsha
DOI:10.1021/acs.jcim.3c01211delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Deep learning methods that predict protein-ligand binding have recently been used for structure-based virtual screening. Many such models have been trained using protein-ligand complexes with known crystal structures and activities from the PDBBind data set. However, because PDBbind only includes 20K complexes, models typically fail to generalize to new targets, and model performance is on par with models trained with only ligand information. Conversely, the ChEMBL database contains a wealth of chemical activity information but includes no information about binding poses. We introduce BigBind, a data set that maps ChEMBL activity data to proteins from the CrossDocked data set. BigBind comprises 583 K ligand activities and includes 3D structures of the protein binding pockets. Additionally, we augmented the data by adding an equal number of putative inactives for each target. Using this data, we developed Banana (basic neural network for binding affinity), a neural network-based model to classify active from inactive compounds, defined by a 10 mu M cutoff. Our model achieved an AUC of 0.72 on BigBind's test set, while a ligand-only model achieved an AUC of 0.59. Furthermore, Banana achieved competitive performance on the LIT-PCBA benchmark (median EF1% 1.81) while running 16,000 times faster than molecular docking with Gnina. We suggest that Banana, as well as other models trained on this data set, will significantly improve the outcomes of prospective virtual screening tasks.
Keyword:
DOCKING

期刊

Journal of Chemical Information and Modeling 封面图
Journal of Chemical Information and Modeling
IF:
5.3
论文数:
9.1K
被引数:
4.0W

机构

U
university of north carolina
学者数:
7.4W
论文数: 6.5W
被引数: 93
U
University of North Carolina Chapel Hill
学者数:
3.9W
论文数: 3.1W
被引数: 46
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
UniProt: the Universal Protein knowledgebaseUniProt: 通用蛋白质知识库
err2004-01-01
err6.0K
errOAAI
errApweiler, R; Bairoch, A; Wu, CH; Barker, WC; Boeckmann, B; Ferro, S; Gasteiger, E; Huang, HZ; Lopez, R; Magrane, M; Martin, MJ; Natale, DA; O'Donovan, C; Redaschi, N; Yeh, LSL
err分享
err收藏
Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules使用数据驱动的分子连续表示的自动化学设计
err2018-01-12
err2.5K
errOAAI
errGomez-Bombarelli, Rafael; Wei, Jennifer N.; Duvenaud, David; Hernandez-Lobato, Jose Miguel; Sanchez-Lengeling, Benjamin; Sheberla, Dennis; Aguilera-Iparraguirre, Jorge; Hirzel, Timothy D.; Adams, Ryan P.; Aspuru-Guzik, Alan
err分享
err收藏
GNINA 1.0: molecular docking with deep learningGNINA 1.0: 分子对接与深度学习
err2021-06-09
err300
errOAAI
errMcNutt, Andrew T.; Francoeur, Paul; Aggarwal, Rishal; Masuda, Tomohide; Meli, Rocco; Ragoza, Matthew; Sunseri, Jocelyn; Koes, David Ryan
err分享
err收藏
Virtual Screening with Gnina 1.0
err2021-12-04
err32
errOAAI
errSunseri, Jocelyn; Koes, David Ryan
err分享
err收藏
学者 查看更多内容