返回
Feature selection with missing data using mutual information estimators
DOI:10.1016/j.neucom.2012.02.031.png)
摘要
En 中文
Feature selection is an important preprocessing task for many machine learning and pattern recognition applications, including regression and classification. Missing data are encountered in many real-world problems and have to be considered in practice. This paper addresses the problem of feature selection in prediction problems where some occurrences of features are missing. To this end, the well-known mutual information criterion is used. More precisely, it is shown how a recently introduced nearest neighbors based mutual information estimator can be extended to handle missing data. This estimator has the advantage over traditional ones that it does not directly estimate any probability density function. Consequently, the mutual information may be reliably estimated even when the dimension of the space increases. Results on artificial as well as real-world datasets indicate that the method is able to select important features without the need for any imputation algorithm, under the assumption of missing completely at random data. Moreover, experiments show that selecting the features before imputing the data generally increases the precision of the prediction models, in particular when the proportion of missing data is high. (C) 2012 Elsevier B.V. All rights reserved.
Keyword:
Feature selection
Missing data
Mutual information
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
The New Targeted Nanoparticles: Tf-Peg-PLL-PLGA Combined with Daunorubicin Overcome Hypoxia Induced Drug Resistance of K562 Cells新型靶向纳米颗粒:Tf-Peg-PLL-PLGA与柔红霉素联合应用克服了K562细胞缺氧诱导的药物抵抗
Blood
IF0

