arrow
返回

Benchmarking missing-values approaches for predictive models on health databases

delete2022-04-15
delete0
delete
OA
AI
A
Alexandre Perez-Lebel *
G
Gaël Varoquaux
M
Marine Le Morvan
J
Julie Josse
J
Jean‐Baptiste Poline
DOI:10.1093/gigascience/giac013delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Background As databases grow larger, it becomes harder to fully control their collection, and they frequently come with missing values. These large databases are well suited to train machine learning models, e.g., for forecasting or to extract biomarkers in biomedical settings. Such predictive approaches can use discriminative-rather than generative-modeling and thus open the door to new missing-values strategies. Yet existing empirical evaluations of strategies to handle missing values have focused on inferential statistics. Results Here we conduct a systematic benchmark of missing-values strategies in predictive models with a focus on large health databases: 4 electronic health record datasets, 1 population brain imaging database, 1 health survey, and 2 intensive care surveys. Using gradient-boosted trees, we compare native support for missing values with simple and state-of-the-art imputation prior to learning. We investigate prediction accuracy and computational time. For prediction after imputation, we find that adding an indicator to express which values have been imputed is important, suggesting that the data are missing not at random. Elaborate missing-values imputation can improve prediction compared to simple strategies but requires longer computational time on large data. Learning trees that model missing values-with missing incorporated attribute-leads to robust, fast, and well-performing predictive modeling. Conclusions Native support for missing values in supervised machine learning predicts better than state-of-the-art imputation with much less computational cost. When using imputation, it is important to add indicator columns expressing which values have been imputed.
Keyword:
missing values
machine learning
supervised learning
benchmark
imputation
multiple imputation
bagging
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

GigaScience 封面图
GigaScience
IF:
3.9
论文数:
1.6K
被引数:
1.2W

机构

U
universite de montpellier
学者数:
3.8W
论文数: 2.6W
被引数: 46
M
McGill University
学者数:
5.5W
论文数: 4.9W
被引数: 7.0W
引用论文

引用论文

HER2 Amplification and Overexpression Is Not Present in Pediatric Osteosarcoma: A Tissue Microarray Study
err2005-09-01
err0
PREAI
errGino R. Somers; Michael Ho; Maria Zielenska; Jeremy A. Squire; Paul S. Thorner
err分享
err收藏
Self-assembled spongy-like MnO2 electrode materials for supercapacitors自组装海绵状MnO2超级电容器电极材料
err2012-08-01
err0
PREAI
errMeng Dong; Yu Xin Zhang; Hong Fang Song; Xin Qiu; Xiao Dong Hao; Chuan Pu Liu; Yuan Yuan; Xin Lu Li; Jia Mu Huang
err分享
err收藏
To the editor
err1993-05-01
err0
PREAI
errB.J. Shannon
err分享
err收藏
学者 查看更多内容