arrow
Return

Missing Value Imputation via Clusterwise Linear Regression

delete2020-01-01
delete33
PRE
AI
N
Napsu Karmitsa *
S
Sona Taheri
A
Adil Bagirov
DOI:10.1109/TKDE.2020.3001694delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In this paper a new method of preprocessing incomplete data is introduced. The method is based on clusterwise linear regression and it combines two well-known approaches for missing value imputation: linear regression and clustering. The idea is to approximate missing values using only those data points that are somewhat similar to the incomplete data point. A similar idea is used also in clustering based imputation methods. Nevertheless, here the linear regression approach is used within each cluster to accurately predict the missing values, and this is done simultaneously to clustering. The proposed method is tested using some synthetic and real-world data sets and compared with other algorithms for missing value imputations. Numerical results demonstrate that this method produces the most accurate imputations in MCAR and MAR data sets with a clear structure and the percentages of missing data no more than 25 percent.
Keywords:
TV
Data analysis
incomplete data
imputation
clusterwise linear regression
nonsmooth optimization
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Knowledge and Data Engineering cover
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
Papers:
6.8K
Citations:
3.2W

Organization

U
University of Turku
Scholars:
1.7W
Papers: 1.5W
Citations: 2.0W
F
Federation University Australia
Scholars:
2.0K
Papers: 2.3K
Citations: 17