Return
Online streaming feature selection for high-dimensional small-sample data
DOI:10.1007/s13042-024-02416-9.png)
Abstract
En 中文
Within the domain of high-dimensional small-sample data classification tasks, there are several significant challenges. The feature space of samples typically has high dimensionality, and the features may be extracted sequentially. In addition, the distribution of data classes is imbalanced. To address these challenges, a novel online feature selection algorithm specifically designed for this scenario is presented in this paper. First, an adaptive neighborhood relation based on class density is proposed, which fully utilizes the distribution of the target sample's class. Second, a neighborhood consistency metric is defined based on the proposed neighborhood relation. Moreover, the proposed online feature selection algorithm consists of three main phases: significance analysis, correlation analysis, and redundancy update. Comprehensive experimental studies on 12 datasets illustrate that our method significantly enhances the prediction of minority class samples, compared to several popular online streaming feature selection algorithms.
Keywords:
Consistent analysis
Class imbalance learning
Feature selection
High-dimensional small-sample data
Journal
IF:
2.7
Papers:
3.1K
Citations:
5.6K

