arrow
Return

Biases in feature selection with missing data

delete2019-05-01
delete16
PRE
AI
B
Borja Seijo-Pardo
A
Amparo Alonso‐Betanzos *
K
Kristin P. Bennett
V
Verónica Bolón‐Canedo
J
Julie Josse
I
Isabelle Guyon
DOI:10.1016/j.neucom.2018.10.085delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Feature selection is of great importance for two possible scenarios: (1) prediction, i.e., improving (or minimally degrading) the predictions of a target variable while discarding redundant or uninformative features and (2) discovery, i.e., identifying features that are truly dependent on the target and may be genuine causes to be determined in experimental verifications (for example for the task of drug target discovery in genomics). In both cases, if variables have a large number of missing values, imputing them may lead to false positives; features that are not associated with the target become dependent as a result of imputation. In the first scenario, this may not harm prediction, but in the second one, it will erroneously select irrelevant features. In this paper, we study the risk/benefit trade-off of missing value imputation in the context of feature selection, using causal graphs to characterize when structural bias arises. Our aim is also to investigate situations in which imputing missing values may be beneficial to reduce false negatives, a situation that might arise when there is a dependency between feature and target, but the dependency is below the significance level when only complete cases are considered. However, the benefits of reducing false negatives must be balanced against the increased number of false positives. In the case of binary target variable and continuous features, the t-test is often used for univariate feature selection. In this paper, we also introduce a de-biased version of the t-test allowing us to reap the benefits of imputation, while not incurring the penalty of increasing the number of false positives. (C) 2019 Elsevier B.V. All rights reserved.
Keywords:
Feature selection
Missing data
De-biased t-test
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

R
rensselaer polytechnic institute
Scholars:
7.0K
Papers: 6.5K
Citations: 6
U
Universidade da Coruna
Scholars:
6.6K
Papers: 5.7K
Citations: 11
M
Microsoft
Scholars:
3.0K
Papers: 2.7K
Citations: 7
E
Ecole Polytechnique
Scholars:
6.5K
Papers: 4.8K
Citations: 211
I
institut polytechnique de paris
Scholars:
1.3W
Papers: 1.0W
Citations: 6
researcher View more organizations