arrow
Return

BigBind: Learning from Nonstructural Data for Structure-Based Virtual Screening

delete2023-12-19
delete3
PRE
AI
M
Michael Brocidiacono *
P
Paul Francoeur
R
Rishal Aggarwal
K
Konstantin I. Popov
D
David Ryan Koes
A
Alexander Tropsha
DOI:10.1021/acs.jcim.3c01211delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep learning methods that predict protein-ligand binding have recently been used for structure-based virtual screening. Many such models have been trained using protein-ligand complexes with known crystal structures and activities from the PDBBind data set. However, because PDBbind only includes 20K complexes, models typically fail to generalize to new targets, and model performance is on par with models trained with only ligand information. Conversely, the ChEMBL database contains a wealth of chemical activity information but includes no information about binding poses. We introduce BigBind, a data set that maps ChEMBL activity data to proteins from the CrossDocked data set. BigBind comprises 583 K ligand activities and includes 3D structures of the protein binding pockets. Additionally, we augmented the data by adding an equal number of putative inactives for each target. Using this data, we developed Banana (basic neural network for binding affinity), a neural network-based model to classify active from inactive compounds, defined by a 10 mu M cutoff. Our model achieved an AUC of 0.72 on BigBind's test set, while a ligand-only model achieved an AUC of 0.59. Furthermore, Banana achieved competitive performance on the LIT-PCBA benchmark (median EF1% 1.81) while running 16,000 times faster than molecular docking with Gnina. We suggest that Banana, as well as other models trained on this data set, will significantly improve the outcomes of prospective virtual screening tasks.
Keywords:
DOCKING

Journal

Journal of Chemical Information and Modeling cover
Journal of Chemical Information and Modeling
IF:
5.3
Papers:
9.1K
Citations:
4.0W

Organization

U
university of north carolina
Scholars:
7.4W
Papers: 6.5W
Citations: 93
U
University of North Carolina Chapel Hill
Scholars:
3.9W
Papers: 3.1W
Citations: 46
Cited Papers

Cited Papers

AutoDock Vina 1.2.0: New Docking Methods, Expanded Force Field, and Python Bindings
err2021-07-19
err2.8K
errOAAI
errEberhardt, Jerome; Santos-Martins, Diogo; Tillack, Andreas F.; Forli, Stefano
errShare
errSave
err
IF0
err
err0
PREAI
err
errShare
errSave
LIT-PCBA: An Unbiased Data Set for Machine Learning and Virtual Screening
err2020-04-13
err127
errOAAI
errTran-Nguyen, Viet-Khoa; Jacquemard, Celien; Rognan, Didier
errShare
errSave
UniProt: the Universal Protein knowledgebase
err2004-01-01
err6.0K
errOAAI
errApweiler, R; Bairoch, A; Wu, CH; Barker, WC; Boeckmann, B; Ferro, S; Gasteiger, E; Huang, HZ; Lopez, R; Magrane, M; Martin, MJ; Natale, DA; O'Donovan, C; Redaschi, N; Yeh, LSL
errShare
errSave
Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules
err2018-01-12
err2.5K
errOAAI
errGomez-Bombarelli, Rafael; Wei, Jennifer N.; Duvenaud, David; Hernandez-Lobato, Jose Miguel; Sanchez-Lengeling, Benjamin; Sheberla, Dennis; Aguilera-Iparraguirre, Jorge; Hirzel, Timothy D.; Adams, Ryan P.; Aspuru-Guzik, Alan
errShare
errSave
GNINA 1.0: molecular docking with deep learning
err2021-06-09
err300
errOAAI
errMcNutt, Andrew T.; Francoeur, Paul; Aggarwal, Rishal; Masuda, Tomohide; Meli, Rocco; Ragoza, Matthew; Sunseri, Jocelyn; Koes, David Ryan
errShare
errSave
Virtual Screening with Gnina 1.0
err2021-12-04
err32
errOAAI
errSunseri, Jocelyn; Koes, David Ryan
errShare
errSave
researcher View more