Return
A complexity-driven resampling approach for imbalanced learning
DOI:10.1016/j.asoc.2026.116355.png)
Abstract
En 中文
Imbalanced datasets pose significant challenges for machine learning (ML) models, which often lead to biased decisions and poor predictions for the minority class. While existing resampling methods address class imbalance, they sometimes neglect data complexity, which can lead to suboptimal performance in complex distributions. To address this, we propose a complexity-driven resampling method using two-stage clustering. For the majority class, K-Means and hierarchical clustering identify and remove extremely hard-to-classify samples based on complexity measures. For the minority class, K-Means clustering is applied first, and then new samples are generated within each cluster based on cluster complexity. The aim of this approach is to balance the datasets while preserving their inherent distribution. By employing different clustering methods at various stages, the method effectively captures clusters of diverse shapes and sizes, thereby enhancing clustering performance on complex datasets. The proposed method is then evaluated using three popular ML models on 27 public datasets from KEEL and four additional datasets from Kaggle using a rigorous 5-fold cross-validation protocol. The results show that the method significantly improves ML classifiers’ performance for imbalanced datasets, achieving the best average F1-Scores and Recall, alongside highly competitive G-Mean and AUC scores, compared to traditional baselines and recent state-of-the-art resampling techniques. This study indicates that dataset complexity significantly affects the performance of most ML models, indicating that careful consideration should be given when choosing resampling strategies for ML models when the dataset becomes more complex. Ultimately, the proposed method provides a highly reliable framework for real-world scenarios characterized by severe class overlap and noisy feature spaces, ensuring that underlying data structures are preserved during the resampling process.
Keywords:
Machine learning
Data complexity
Imbalanced data
Instance-level complexity
Data balancing
Journal
IF:
6.6
Papers:
1.4W
Citations:
4.8W
Organization
No organization information available
Cited Papers
No cited papers available

