arrow
返回

Mixed Data Imputation Using Generative Adversarial Networks

delete2022-01-01
delete7
delete
OA
AI
W
Wasif Khan
N
Nazar Zaki *
A
Amir Ahmad
M
Mohammad Mehedy Masud
L
Luqman Ali
N
Nasloon Ali
L
Luai A. Ahmed
DOI:10.1109/ACCESS.2022.3218067delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Missing values are common in real-world datasets and pose a significant challenge to the performance of statistical and machine learning models. Generally, missing values are imputed using statistical methods, such as the mean, median, mode, or machine learning approaches. These approaches are limited to either numerical or categorical data. Imputation in mixed datasets that contain both numerical and categorical attributes is challenging and has received little attention. Machine learning-based imputation algorithms usually require a large amount of training data. However, obtaining such data is difficult. Furthermore, no considerate work has been conducted in the literature that focuses on the effects of the training and testing size with increasing amounts of missing data. To address this gap, we proposed that increasing the amount of training data will improve imputation performance. We first used generative adversarial network (GAN) methods to increase the amount of training data. We considered two state-of-the-art GANs (tabular and conditional tabular) to add synthetic samples using observed data with different synthetic sample ratios. We then used three state-of-the-art imputation models that can handle mixed data: MissForest, multivariate imputation by chained equations, and denoising auto encoder (DAE). We proposed robust experimental setups on four publicly available datasets with different training-testing data divisions that have increasing missingness ratios. Extensive experimental results show that incorporating synthetic samples with training data achieves better performance compared to the baseline methods for mixed data imputation in both categorical and numerical variables, especially for large missingness ratios.
Keyword:
Training data
Generative adversarial networks
Generators
Statistics
Data models
Machine learning algorithms
Prediction algorithms
Sequential analysis
Multivariate regression
Mixed data imputation
missing data
GANs
miss forest
MICE
denoising auto encoders

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

U
United Arab Emirates University
学者数:
8.8K
论文数: 7.4K
被引数: 10.0K
引用论文

引用论文

Survey of State-of-the-Art Mixed Data Clustering Algorithms
err2019-01-01
err143
errOAAI
errAhmad, Amir; Khan, Shehroz S.
err分享
err收藏
err分享
err收藏
err分享
err收藏
Evolution of abstracts presented at the annual scientific meetings of academic emergency medicine
err1999-10-01
err0
PREAI
errAdam J Singer; Clark S Homan; Michael Brody; Henry C Thode; Judd E Hollander
err分享
err收藏
A Survey on Data Imputation Techniques: Water Distribution System as a Use Case
err2018-01-01
err68
errOAAI
errOsman, Muhammad S.; Abu-Mahfouz, Adnan M.; Page, Philip R.
err分享
err收藏
学者 查看更多内容