Return
Undersampling method based on minority class density for imbalanced data
DOI:10.1016/j.eswa.2024.123328.png)
Abstract
En 中文
Imbalanced data severely hinder the classification performance of learning-based algorithms and attract a great deal of attention from researchers. The undersampling method of selecting the majority class from samples through the prototype is one of the most common techniques to address class imbalance. However, traditional undersampling-based methods inherently lead to information loss. Meanwhile, they tend to perform undersampling for the majority class using distance-based locality methods, resulting in completely ignoring the density of the minority class samples. To this end, this paper proposes a novel undersampling method that incorporates the density distribution information of the minority class. Specifically, the probability density distribution of the minority class samples is learned by kernel density estimation, and the majority class samples located in the high-density band of the minority class are removed through filtering. Based on this, sampling fitness is proposed to evaluate the desirable value of each majority class sample to select informationrich samples. The research was carried out in 25 publicly available datasets and compared with state -of -art methods. The results show that our approach offers clear superiority.
Keywords:
Undersampling
Imbalanced data
Kernel density estimation
Classification
Data mining
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W

