返回
Deep learning versus conventional methods for missing data imputation: A review and comparative study
DOI:10.1016/j.eswa.2023.120201.png)
摘要
En 中文
Deep learning models have been recently proposed in the applications of missing data imputation. In this paper, we review the popular statistical, machine learning, and deep learning approaches, and discuss the advantages and disadvantages of these methods. We conduct a comprehensive numerical study to compare the performance of several widely-used imputation methods for incomplete tabular (structured) data. Specifically, we compare the deep learning methods: generative adversarial imputation networks (GAIN) with onehot encoding, GAIN with embedding, variational auto-encoder (VAE) with onehot encoding, and VAE with embedding versus two conventional methods: multiple imputation by chained equations (MICE) and missForest. Seven real benchmark datasets and three simulated datasets are considered, including various scenarios with different feature types under different levels of sample sizes. The missing data are generated based on different missing ratios and three kinds of missing mechanisms: missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR). Our experiments show that, for small or moderate sample sizes, the conventional methods establish better robustness and imputation performance than the deep learning methods. GAINs only perform well in the case of MCAR and often fail in the cases of MAR and MNAR. VAEs are easy to fall into mode collapse in all missing mechanisms. We conclude that the conventional methods, MICE and missForest, are preferable for practitioners to deal with missing data imputation for tabular data with a limited sample size (i.e., n < 30, 000) in real case analyses.
Keyword:
Missing data imputation
Deep learning
Generative networks
MICE
MissForest
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W
机构
引用论文
Comparison of Random Forest and Parametric Imputation Models for Imputing Missing Data Using MICE: A CALIBER Study随机森林和参数填补模型在小鼠缺失数据填补中的比较: 一项口径研究
Personality Differences among Patients with Chronic Aphasia Predict Improvement in Speech-Language Therapy慢性失语症患者的人格差异预示着言语语言治疗的改善
Generative adversarial networks for imputing missing data for big data clinical research用于大数据临床研究的缺失数据的生成对抗网络
Embedded Data Imputation for Environmental Intelligent Sensing: A Case Study环境智能感知的嵌入式数据插补: 案例研究
SENSORS
IF3.5

