arrow
Return

Online streaming feature selection for high-dimensional small-sample data

delete2024-10-16
delete0
PRE
AI
K
Kuangfeng Gong *
G
Guohe Li
L
Lingyun Guo
Y
Yaojin Lin
DOI:10.1007/s13042-024-02416-9delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Within the domain of high-dimensional small-sample data classification tasks, there are several significant challenges. The feature space of samples typically has high dimensionality, and the features may be extracted sequentially. In addition, the distribution of data classes is imbalanced. To address these challenges, a novel online feature selection algorithm specifically designed for this scenario is presented in this paper. First, an adaptive neighborhood relation based on class density is proposed, which fully utilizes the distribution of the target sample's class. Second, a neighborhood consistency metric is defined based on the proposed neighborhood relation. Moreover, the proposed online feature selection algorithm consists of three main phases: significance analysis, correlation analysis, and redundancy update. Comprehensive experimental studies on 12 datasets illustrate that our method significantly enhances the prediction of minority class samples, compared to several popular online streaming feature selection algorithms.
Keywords:
Consistent analysis
Class imbalance learning
Feature selection
High-dimensional small-sample data

Journal

International Journal of Machine Learning and Cybernetics cover
International Journal of Machine Learning and Cybernetics
IF:
2.7
Papers:
3.1K
Citations:
5.6K

Organization

M
Minnan Normal University
Scholars:
2.1K
Papers: 1.3K
Citations: 0
C
china university of petroleum
Scholars:
4.1W
Papers: 2.7W
Citations: 30