返回
Missing Value Imputation via Clusterwise Linear Regression
DOI:10.1109/TKDE.2020.3001694.png)
摘要
En 中文
In this paper a new method of preprocessing incomplete data is introduced. The method is based on clusterwise linear regression and it combines two well-known approaches for missing value imputation: linear regression and clustering. The idea is to approximate missing values using only those data points that are somewhat similar to the incomplete data point. A similar idea is used also in clustering based imputation methods. Nevertheless, here the linear regression approach is used within each cluster to accurately predict the missing values, and this is done simultaneously to clustering. The proposed method is tested using some synthetic and real-world data sets and compared with other algorithms for missing value imputations. Numerical results demonstrate that this method produces the most accurate imputations in MCAR and MAR data sets with a clear structure and the percentages of missing data no more than 25 percent.
Keyword:
TV
Data analysis
incomplete data
imputation
clusterwise linear regression
nonsmooth optimization
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
10.4
论文数:
6.8K
被引数:
3.2W
机构
引用论文
Integrated Optimization Design of Combined Cooling, Heating, and Power System Coupled with Solar and Biomass Energy
Energies
IF0

