arrow
Return

A practical guide to machine-learning scoring for structure-based virtual screening

delete2023-10-16
delete23
PRE
AI
V
Viet‐Khoa Tran‐Nguyen
M
Muhammad Junaid
S
Saw Simeon
P
Pedro J. Ballester *
DOI:10.1038/s41596-023-00885-wdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Structure-based virtual screening (SBVS) via docking has been used to discover active molecules for a range of therapeutic targets. Chemical and protein data sets that contain integrated bioactivity information have increased both in number and in size. Artificial intelligence and, more concretely, its machine-learning (ML) branch, including deep learning, have effectively exploited these data sets to build scoring functions (SFs) for SBVS against targets with an atomic-resolution 3D model (e.g., generated by X-ray crystallography or predicted by AlphaFold2). Often outperforming their generic and non-ML counterparts, target-specific ML-based SFs represent the state of the art for SBVS. Here, we present a comprehensive and user-friendly protocol to build and rigorously evaluate these new SFs for SBVS. This protocol is organized into four sections: (i) using a public benchmark of a given target to evaluate an existing generic SF; (ii) preparing experimental data for a target from public repositories; (iii) partitioning data into a training set and a test set for subsequent target-specific ML modeling; and (iv) generating and evaluating target-specific ML SFs by using the prepared training-test partitions. All necessary code and input/output data related to three example targets (acetylcholinesterase, HMG-CoA reductase, and peroxisome proliferator-activated receptor-alpha) are available at https://github. com/vktrannguyen/MLSF-protocol, can be run by using a single computer within 1 week and make use of easily accessible software/programs (e.g., Smina, CNN-Score, RF-Score-VS and DeepCoy) and web resources. Our aim is to provide practical guidance on how to augment training data to enhance SBVS performance, how to identify the most suitable supervised learning algorithm for a data set, and how to build an SF with the highest likelihood of discovering target-active molecules within a given compound library.
Keywords:
ASSAY INTERFERENCE COMPOUNDS
LIGAND BINDING-AFFINITY
SWISS-MODEL REPOSITORY
APPLICABILITY DOMAIN
MOLECULAR DOCKING
COMPOUNDS PAINS
DATA SETS
PROTEIN
DISCOVERY
ACCURACY

Journal

Nature Protocols cover
Nature Protocols
IF:
16
Papers:
4.0K
Citations:
5.6W

Organization

A
aix-marseille universite
Scholars:
3.8W
Papers: 2.7W
Citations: 77
Cited Papers

Cited Papers

errShare
errSave
DNA yield and quality of saliva samples and suitability for large-scale epidemiological studies in children
err2011-04-12
err0
PREAI
errA C Koni; R A Scott; G Wang; M E S Bailey; J Peplies; K Bammann; Y P Pitsiladis
errShare
errSave
Bias-motivated bullying and psychosocial problems: Implications for HIV risk behaviors among young men who have sex with men
err2013-06-25
err0
PREAI
errMichael Jonathan Li; Anthony DiStefano; Michele Mouttapa; Jasmeet K. Gill
errShare
errSave
Molecular persistent spectral image (Mol-PSI) representation for machine learning models in drug design
err2021-12-28
err17
PREAI
errJiang, Peiran; Chi, Ying; Li, Xiao-Shuang; Liu, Xiang; Hua, Xian-Sheng; Xia, Kelin
errShare
errSave
Constructing and Validating High-Performance MIEC-SVM Models in Virtual Screening for Kinases: A Better Way for Actives Discovery
err2016-04-22
err63
errOAAI
errSun, Huiyong; Pan, Peichen; Tian, Sheng; Xu, Lei; Kong, Xiaotian; Li, Youyong; Li, Dan; Hou, Tingjun
errShare
errSave
Drugs for bad bugs: confronting the challenges of antibacterial discovery
err2006-12-08
err2.4K
PREAI
errPayne, David J.; Gwynn, Michael N.; Holmes, David J.; Pompliano, David L.
errShare
errSave
errShare
errSave
researcher View more