Return
Grouping-based Oversampling in Kernel Space for Imbalanced Data Classification
DOI:10.1016/j.patcog.2022.108992.png)
Abstract
En 中文
The class-imbalanced classification is a difficult problem because not only traditional classifiers are more biased towards the majority classes and inclined to generate incorrect predictions, but also the existing algorithms often have difficulty tackling this kind of problem with the class overlapping. Oversampling is a widely used and effective method to obtain balanced samples for imbalanced data, but the existing oversampling methods usually result in more serious class overlapping due to improper choice of the ref-erence samples. To circumvent this shortcoming, according to the different possibilities of minority class samples appearing in the overlapping regions in the feature space, a grouping scheme for the minor-ity class samples is first designed to identify the overlapping region samples. Then, a new oversampling method based on this grouping scheme is proposed to make the new samples far away from the over-lapping region and rectify the decision boundary properly. Subsequently, a new effective classification algorithm is developed for imbalanced data. Extensive experiments show that the proposed algorithm is superior to the seventeen benchmark algorithms in terms of three performance metrics, especially on high imbalance ratio data sets. (c) 2022 Elsevier Ltd. All rights reserved.
Keywords:
Imbalanced data classification
Kernel method
Support vector machine
Oversampling
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W

