返回
Imbalanced Data Problem in Machine Learning: A Review
DOI:10.1109/ACCESS.2025.3531662.png)
摘要
En 中文
One of the prominent challenges encountered in real-world data is an imbalance, characterized by unequal distribution of observations across different target classes, which complicates achieving accurate model classifications. This survey delves into various machine learning techniques developed to address the difficulties posed by imbalanced data. It discusses data-level methods such as oversampling and undersampling, algorithm-level solutions including ensemble learning and specific algorithm adjustments, cost-sensitive algorithms, and hybrid strategies that combine multiple approaches. Moreover, this paper emphasizes the crucial role of evaluation methods like Precision, F1 Score, Recall, G-mean, and AUC in measuring the effectiveness of these strategies under imbalanced conditions. A detailed review of recent research articles helps pinpoint persistent gaps in generalizability, scalability, and robustness across these methods, underscoring the necessity for ongoing improvements. The survey seeks to offer an extensive overview of current approaches that improve the efficiency and effectiveness of machine learning models dealing with imbalanced datasets, thus equipping researchers with the insights needed to develop robust and effective models ready for real-world application.
Keyword:
Data models
Machine learning
Classification algorithms
Machine learning algorithms
Training
Reviews
Fraud
Surveys
Ensemble learning
Data augmentation
Imbalanced data
machine learning
balance techniques
evaluation methods
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
A Novel Imbalanced Ensemble Learning in Software Defect Predication一种新的软件缺陷预测中的不平衡集成学习
IEEE ACCESS
IF3.6
Arsenic removal from aqueous solutions by adsorption using novel MIL-53(Fe) as a highly efficient adsorbent使用新型MIL-53(Fe) 作为高效吸附剂通过吸附从水溶液中去除砷
RSC Advances
IF0

