arrow
返回

A Framework for Effective Application of Machine Learning to Microbiome-Based Classification Problems

delete2020-06-30
delete116
delete
OA
AI
B
Begüm D. Topçuoğlu
N
Nicholas A. Lesniak
M
Mack T. Ruffin
J
Jenna Wiens
P
Patrick D. Schloss *
DOI:10.1128/mBio.00434-20delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Machine learning (ML) modeling of the human microbiome has the potential to identify microbial biomarkers and aid in the diagnosis of many diseases such as inflammatory bowel disease, diabetes, and colorectal cancer. Progress has been made toward developing ML models that predict health outcomes using bacterial abundances, but inconsistent adoption of training and evaluation methods call the validity of these models into question. Furthermore, there appears to be a preference by many researchers to favor increased model complexity over interpretability. To overcome these challenges, we trained seven models that used fecal 16S rRNA sequence data to predict the presence of colonic screen relevant neoplasias (SRNs) (n = 490 patients, 261 controls and 229 cases). We developed a reusable open-source pipeline to train, validate, and interpret ML models. To show the effect of model selection, we assessed the predictive performance, interpretability, and training time of L2-regularized logistic regression, L1- and L2-regularized support vector machines (SVM) with linear and radial basis function kernels, a decision tree, random forest, and gradient boosted trees (XGBoost). The random forest model performed best at detecting SRNs with an area under the receiver operating characteristic curve (AUROC) of 0.695 (interquartile range [IQR], 0.651 to 0.739) but was slow to train (83.2 h) and not inherently interpretable. Despite its simplicity, L2-regularized logistic regression followed random forest in predictive performance with an AUROC of 0.680 (IQR, 0.625 to 0.735), trained faster (12 min), and was inherently interpretable. Our analysis highlights the importance of choosing an ML approach based on the goal of the study, as the choice will inform expectations of performance and interpretability. IMPORTANCE Diagnosing diseases using machine learning (ML) is rapidly being adopted in microbiome studies. However, the estimated performance associated with these models is likely overoptimistic. Moreover, there is a trend toward using black box models without a discussion of the difficulty of interpreting such models when trying to identify microbial biomarkers of disease. This work represents a step toward developing more-reproducible ML practices in applying ML to microbiome research. We implement a rigorous pipeline and emphasize the importance of selecting ML models that reflect the goal of the study. These concepts are not particular to the study of human health but can also be applied to environmental microbiology studies.
Keyword:
16S rRNA gene
colon cancer
machine learning
microbial ecology
microbiome
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

mBio 封面图
mBio
IF:
4.7
论文数:
8.6K
被引数:
3.7W

机构

U
University of Michigan
学者数:
6.4W
论文数: 5.3W
被引数: 124
U
university of michigan system
学者数:
9.1W
论文数: 8.6W
被引数: 133
引用论文

引用论文

Effect of feeding regimen on the fatty acid profile of sheep bulk tank milk
err2018-08-08
err0
PREAI
errErica Renes; Fernando De la Fuente; Domingo Fernández; María Eugenia Tornadijo; José María Fresno
err分享
err收藏
Pretreatment gut microbiome predicts chemotherapy-related bloodstream infection
err2016-04-28
err156
errOAAI
errMontassier, Emmanuel; Al-Ghalith, Gabriel A.; Ward, Tonya; Corvec, Stephane; Gastinne, Thomas; Potel, Gilles; Moreau, Phillipe; de la Cochetiere, Marie France; Batard, Eric; Knights, Dan
err分享
err收藏
Looking for a Signal in the Noise: Revisiting Obesity and the Microbiome
errMBIO
IF4.7
err2016-09-07
err460
errOAAI
errSze, Marc A.; Schloss, Patrick D.
err分享
err收藏
学者 查看更多内容