返回
Missing value imputation on missing completely at random data using multilayer perceptrons
DOI:10.1016/j.neunet.2010.09.008.png)
摘要
En 中文
Data mining is based on data files which usually contain errors in the form of missing values. This paper focuses on a methodological framework for the development of an automated data imputation model based on artificial neural networks. Fifteen real and simulated data sets are exposed to a perturbation experiment, based on the random generation of missing values. These data set sizes range from 47 to 1389 records. A perturbation experiment was performed for each data set where the probability of missing value was set to 0.05. Several architectures and learning algorithms for the multilayer perceptron are tested and compared with three classic imputation procedures: mean/mode imputation, regression and hot-deck. The obtained results, considering different performance measures, not only suggest this approach improves the quality of a database with missing values, but also the best results are clearly obtained using the Multilayer Perceptron model in data sets with categorical variables. Three learning rules (Levenberg-Marquardt, BFGS Quasi-Newton and Conjugate Gradient Fletcher-Reeves Update) and a small number of hidden nodes are recommended. (C) 2010 Elsevier Ltd. All rights reserved.
Keyword:
Multilayer perceptron
Hot-deck model
Imputation
Mean/mode model
Missing data
Regression model
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.3
论文数:
7.8K
被引数:
3.0W
机构
引用论文
Quantitative DNA variation between and within chromosome complements of Vigna species (Fabaceae)
Genetica
IF0
New Species and Combinations in Vigna Subgenus Ceratotropis (Piper) Verdc. (Leguminosae, Phaseoleae)
Kew Bulletin
IF0
没有更多内容

