arrow
Return

Proportional clustering-based undersampling for imbalanced data classification

delete2025-09-29
delete0
PRE
AI
C
Chengshuo Zhang
Z
Zhanrong Shi
W
Wangwei Lu
赵津 cover
赵津 (Jin Zhao)
冯朔 cover
冯朔 (Shuo Feng) *
M
Mingliang Xu
DOI:10.1007/s10115-025-02593-1delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Class imbalance is an important challenge in machine learning and data mining, as it hinders the detection of rare but important instances. Clustering-based undersampling methods are widely used to address this issue. However, they often struggle to choose appropriate clustering algorithms and identify representative instances, resulting in suboptimal resampling. In this paper, a clustering-based undersampling method is proposed to address the class imbalance problem. The method uses the DBSCAN algorithm for clustering. The number of instances to select from each cluster is determined proportionally, and a linear model optimized by Differential Evolution is used to identify the specific instances to retain, completing the undersampling process. Comparative experiments are conducted on 30 datasets using five classifiers. The results demonstrate that the proposed method significantly outperforms baseline methods in terms of MCC, F-measure, and AUC. Additional experiments further show the impact of clustering algorithms on resampling performance and highlight the effectiveness of the learning-to-rank algorithm in selecting representative instances.
Keywords:
Class imbalance
Undersampling
Oversampling
Clustering
Learning-to-rank

Journal

Knowledge and Information Systems cover
Knowledge and Information Systems
IF:
3.1
Papers:
517
Citations:
5.2K

Organization

S