返回
A membership-based resampling and cleaning algorithm for multi-class imbalanced overlapping data
DOI:10.1016/j.eswa.2023.122565.png)
摘要
En 中文
Real-world datasets frequently have an imbalanced class distribution, which significantly degrades classifica-tion performance. However, some studies have suggested that the adverse effects of class imbalance occur only when datasets have other intrinsic characteristics (such as class overlap, noise, and data scarcity). Noise and class overlap have the greatest effect. To deal with other intrinsic characteristics that affect classification performance in a multi-class environment, such as class overlap, noise, and data scarcity, we propose a method that can directly handle multi-class overlapping data, called the membership-based multi-class resampling and cleaning (MC-MBRC) algorithm. The proposed method divides samples into safe, overlapping, and noisy areas based on their membership degrees, and then according to the influence of the samples in each area on classification performance, it performs various operations such as noise removal, interpolation oversampling, and energy-based cleaning of the overlapping region. An extensive comparison using various datasets shows that, compared with state-of-the-art methods, the proposed method makes significant statistical improvements in different classification performance metrics and is robust to data containing class overlap and label noise. Furthermore, for datasets that do not have complex intrinsic features, MC-MBRC will not significantly degrade classification performance.
Keyword:
Class overlap
Multi-class imbalance
Oversampling
Resampling
期刊
IF:
7.5
论文数:
2.9W
被引数:
10.2W
机构
引用论文
A dynamic over-sampling procedure based on sensitivity for multi-class problems
PATTERN RECOGNITION
IF7.6
Training cost-sensitive neural networks with methods addressing the class imbalance problem用解决类不平衡问题的方法训练代价敏感的神经网络
SVDD boundary and DPC clustering technique-based oversampling approach for handling imbalanced and overlapped data基于SVDD边界和DPC聚类技术的过采样方法,用于处理不平衡和重叠数据

