返回
Loyalty-SMOTE: Data Synthesis Algorithm for Effective Imbalanced Data Classification
DOI:10.1016/j.neunet.2026.108677.png)
摘要
En 中文
Imbalanced datasets are always problematic in training machine learning models, such that classifiers often struggle to achieve satisfactory performance. Numerous approaches have been developed to tackle imbalanced data problems. Among them, some data-level methods perform linear interpolations between neighboring minority class samples to generate new data points, while others focus on oversampling boundary samples which are specific to certain classes. However, many methods fail to consider scenarios involving noise susceptibility. In this paper, we propose a novel data-level method called the Loyalty-SMOTE algorithm. We introduce the concept of Loyalty to identify noise and boundaries within datasets. After identifying potential noisy datapoints, SMOTE (Synthetic Minority Oversampling Technique) algorithm is applied to oversample the minority class boundary data. Subsequently, a denoising process based on Loyalty is conducted to obtain a balanced dataset. To extend our design, the concept of Attraction is introduced to generalize the denoising technique for multiclass problems. In our study, the SVM (Support Vector Machine) classifier is used as our base learner, extensive experiments are performed to evaluate and compare different algorithms. Our results demonstrate that Loyalty-SMOTE achieved superior performance across multiple metrics on both binary and multiclass UCI datasets. For 30 binary datasets, it achieved the highest scores in 26 datasets (87%) for F1-score, 29 datasets (97%) for AUROC, 26 datasets (87%) for recall, and 27 datasets (90%) for G-mean. For 5 multiclass datasets, our design achieved scores of 0.8317, 0.6153, 0.8537, and 0.6717, respectively.
期刊
IF:
6.3
论文数:
7.8K
被引数:
3.0W
机构
引用论文
Modifying the SMOTE and Safe-Level SMOTE Oversampling Method to Improve PerformanceDolo, K.M.; Mnkandla, E. 修改SMOTE和Safe-Level SMOTE过采样方法以提高性能。载于《数据工程与通信技术讲义》;施普林格自然出版社:新加坡,2022年;第47-59页。 [谷歌学术] [CrossRef]
SMOTE vs. SMOTEENN: A Study on the Performance of Resampling Algorithms for Addressing Class Imbalance in Regression ModelsHusain, G.; Nasef, D.; Jose, R.; Mayer, J.; Bekbolatova, M.; Devine, T.; Toma, M. SMOTE与SMOTEENN:针对回归模型中类别不平衡问题的重采样算法性能研究。Algorithms 2025, 18, 37. [Google Scholar] [CrossRef]
Algorithms
IF0
A Boundary-Information-Based Oversampling Approach to Improve Learning Performance for Imbalanced Datasets
Entropy
IF0

