Return
Model-free feature screening for ultrahigh dimensional data with responses missing not at random
DOI:10.1016/j.jmva.2026.105605.png)
Abstract
En 中文
Feature screening method is an important tool for screening active features in ultrahigh dimensional data analysis. Existing feature screening methods mainly focus on the fully observed data or missing responses at random. But in many applied fields such as biomedicine, social science and epidemiological studies, responses might be subject to nonignorable missingness due to various reasons such as dropout. To this end, this paper proposes a new adjusted Spearman rank correlation to screen active features by incorporating the Spearman rank correlation and its conditional expectation in the presence of nonignorable missing responses. To circumvent the notorious identification problem, we introduce instrumental variables into the propensity score (PS) function, which is specified by a more general semiparametric regression model. A nonparametric imputation method is developed to estimate the adjusted Spearman rank correlation. The proposed method has several desirable merits. First, it is model-free. Second, it is robust to outliers, heavy tailed data and the misspecification of the PS function. Third, under some weaker regularity conditions than existing missing data literature, it has sure screening property and ranking consistency, and can well control the false discovery rate regardless of known or consistently estimated parameters in the PS function. Simulation studies and two real examples are used to investigate the performance of the proposed methodologies.
Keywords:
Feature screening
Instrumental variable
Missing not at random
Nonparametric imputation
Ultrahigh dimensional data
Journal
J
IF:
1.7
Papers:
97
Citations:
5.8K

