arrow
返回

Extrapolation is not the same as interpolation

delete2024-07-23
delete1
delete
OA
AI
Y
Yuxuan Wang *
R
Ross D. King
DOI:10.1007/s10994-024-06591-2delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
We propose a new machine learning formulation designed specifically for extrapolation. The textbook way to apply machine learning to drug design is to learn a univariate function that when a drug (structure) is input, the function outputs a real number (the activity): f(drug) ->\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\rightarrow$$\end{document} activity. However, experience in real-world drug design suggests that this formulation of the drug design problem is not quite correct. Specifically, what one is really interested in is extrapolation: predicting the activity of new drugs with higher activity than any existing ones. Our new formulation for extrapolation is based on learning a bivariate function that predicts the difference in activities of two drugs F(drug1, drug2) ->\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\rightarrow$$\end{document} difference in activity, followed by the use of ranking algorithms. This formulation is general and agnostic, suitable for finding samples with target values beyond the target value range of the training set. We applied the formulation to work with support vector machines , random forests , and Gradient Boosting Machines . We compared the formulation with standard regression on thousands of drug design datasets, gene expression datasets and material property datasets. The test set extrapolation metric was the identification of examples with greater values than the training set, and top-performing examples (within the top 10% of the whole dataset). On this metric our pairwise formulation vastly outperformed standard regression. Its proposed variations also showed a consistent outperformance. Its application in the stock selection problem further confirmed the advantage of this pairwise formulation.
Keyword:
Machine learning
Ranking
Extrapolation
Drug discovery

期刊

Machine Learning 封面图
Machine Learning
IF:
2.9
论文数:
2.7K
被引数:
3.4W

机构

U
University of Cambridge
学者数:
7.7W
论文数: 7.1W
被引数: 13.7W
引用论文

引用论文

HOXA9 is a novel myopia risk gene
err2019-01-23
err0
errOAAI
errChung-Ling Liang; Po-Yuan Hsu; Cheryl S. Ngo; Wei Jie Seow; Neerja Karnani; Hong Pan; Seang-Mei Saw; Suh-Hang H. Juo
err分享
err收藏
ChEMBL: towards direct deposition of bioassay dataChEMBL: 走向生物测定数据的直接沉积
err2018-11-06
err1.3K
errOAAI
errMendez, David; Gaulton, Anna; Bento, A. Patricia; Chambers, Jon; De Veij, Marleen; Felix, Eloy; Magarinos, Maria Paula; Mosquera, Juan F.; Mutowo, Prudence; Nowotka, Michal; Gordillo-Maranon, Maria; Hunter, Fiona; Junco, Laura; Mugumbate, Grace; Rodriguez-Lopez, Milagros; Atkinson, Francis; Bosc, Nicolas; Radoux, ChrisJ; Segura-Cabrera, Aldo; Hersey, Anne; Leach, Andrew R.
err分享
err收藏
When drug discovery meets web search: Learning to Rank for ligand-based virtual screening
err2015-02-13
err29
errOAAI
errZhang, Wei; Ji, Lijuan; Chen, Yanan; Tang, Kailin; Wang, Haiping; Zhu, Ruixin; Jia, Wei; Cao, Zhiwei; Liu, Qi
err分享
err收藏
err分享
err收藏
StructRank: A New Approach for Ligand-Based Virtual Screening
err2010-12-17
err27
PREAI
errRathke, Fabian; Hansen, Katja; Brefeld, Ulf; Mueller, Klaus-Robert
err分享
err收藏
学者 查看更多内容