arrow
返回

A practical guide to machine-learning scoring for structure-based virtual screening

delete2023-10-16
delete23
PRE
AI
V
Viet‐Khoa Tran‐Nguyen
M
Muhammad Junaid
S
Saw Simeon
P
Pedro J. Ballester *
DOI:10.1038/s41596-023-00885-wdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Structure-based virtual screening (SBVS) via docking has been used to discover active molecules for a range of therapeutic targets. Chemical and protein data sets that contain integrated bioactivity information have increased both in number and in size. Artificial intelligence and, more concretely, its machine-learning (ML) branch, including deep learning, have effectively exploited these data sets to build scoring functions (SFs) for SBVS against targets with an atomic-resolution 3D model (e.g., generated by X-ray crystallography or predicted by AlphaFold2). Often outperforming their generic and non-ML counterparts, target-specific ML-based SFs represent the state of the art for SBVS. Here, we present a comprehensive and user-friendly protocol to build and rigorously evaluate these new SFs for SBVS. This protocol is organized into four sections: (i) using a public benchmark of a given target to evaluate an existing generic SF; (ii) preparing experimental data for a target from public repositories; (iii) partitioning data into a training set and a test set for subsequent target-specific ML modeling; and (iv) generating and evaluating target-specific ML SFs by using the prepared training-test partitions. All necessary code and input/output data related to three example targets (acetylcholinesterase, HMG-CoA reductase, and peroxisome proliferator-activated receptor-alpha) are available at https://github. com/vktrannguyen/MLSF-protocol, can be run by using a single computer within 1 week and make use of easily accessible software/programs (e.g., Smina, CNN-Score, RF-Score-VS and DeepCoy) and web resources. Our aim is to provide practical guidance on how to augment training data to enhance SBVS performance, how to identify the most suitable supervised learning algorithm for a data set, and how to build an SF with the highest likelihood of discovering target-active molecules within a given compound library.
Keyword:
ASSAY INTERFERENCE COMPOUNDS
LIGAND BINDING-AFFINITY
SWISS-MODEL REPOSITORY
APPLICABILITY DOMAIN
MOLECULAR DOCKING
COMPOUNDS PAINS
DATA SETS
PROTEIN
DISCOVERY
ACCURACY

期刊

Nature Protocols 封面图
Nature Protocols
IF:
16
论文数:
4.0K
被引数:
5.6W

机构

A
aix-marseille universite
学者数:
3.8W
论文数: 2.7W
被引数: 77
I
institut national de la sante et de la recherche medicale (inserm)
学者数:
11.5W
论文数: 7.5W
被引数: 117
引用论文

引用论文

err分享
err收藏
DNA yield and quality of saliva samples and suitability for large-scale epidemiological studies in children
err2011-04-12
err0
PREAI
errA C Koni; R A Scott; G Wang; M E S Bailey; J Peplies; K Bammann; Y P Pitsiladis
err分享
err收藏
Bias-motivated bullying and psychosocial problems: Implications for HIV risk behaviors among young men who have sex with men
err2013-06-25
err0
PREAI
errMichael Jonathan Li; Anthony DiStefano; Michele Mouttapa; Jasmeet K. Gill
err分享
err收藏
Molecular persistent spectral image (Mol-PSI) representation for machine learning models in drug design
err2021-12-28
err17
PREAI
errJiang, Peiran; Chi, Ying; Li, Xiao-Shuang; Liu, Xiang; Hua, Xian-Sheng; Xia, Kelin
err分享
err收藏
Constructing and Validating High-Performance MIEC-SVM Models in Virtual Screening for Kinases: A Better Way for Actives Discovery
err2016-04-22
err63
errOAAI
errSun, Huiyong; Pan, Peichen; Tian, Sheng; Xu, Lei; Kong, Xiaotian; Li, Youyong; Li, Dan; Hou, Tingjun
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容