返回
A novel Random Forest integrated model for imbalanced data classification problem
DOI:10.1016/j.knosys.2022.109050.png)
摘要
En 中文
In recent years, most researchers focused on the classification problems of imbalanced data sets, and these problems are widely distributed in industrial production and medical research fields. For these highly imbalanced data sets, the ensemble method based on over-sampling is one of the most competitive techniques in the present research. However, the incorrect sampling strategy easily affected the model performance, which increased the training complexity and caused an over-fitting problem. This article proposed an equilibrium ensemble method (DCI-ISSA) with two novel techniques to conquer these shortcomings. Firstly, this paper raised an over-sampling approach (Data Center Interpolation DCI) to offer a counterbalanced data set for the single learner, which can prevent the base learners from the impact of class imbalance. Additionally, we provided a parameter optimization method for Random Forest (RF), which used the Improved Sparrow Search Algorithm (ISSA) to find the optimal parameters for different imbalanced data sets dynamically. These parameters can improve the classification performance of base classifiers and adjust to all kinds of lopsided data sets with distinct sizes. Experimental results showed that the DCI-ISSA-RF model outperforms other famous approaches for the imbalanced data sets with various dimensions.(c) 2022 Published by Elsevier B.V.
Keyword:
Imbalanced data classification
Random Forest
Sparrow Search Algorithm
Oversampling
期刊
K
IF:
7.6
论文数:
1.2W
被引数:
4.5W
机构
暂无机构信息
引用论文
Using Cost-Sensitive Learning and Feature Selection Algorithms to Improve the Performance of Imbalanced Classification使用代价敏感学习和特征选择算法来提高不平衡分类的性能
IEEE ACCESS
IF3.6
Arsenic removal from aqueous solutions by adsorption using novel MIL-53(Fe) as a highly efficient adsorbent使用新型MIL-53(Fe) 作为高效吸附剂通过吸附从水溶液中去除砷
RSC Advances
IF0
An automatic sampling ratio detection method based on genetic algorithm for imbalanced data classification基于遗传算法的不平衡数据分类采样率自动检测方法
An Improving Majority Weighted Minority Oversampling Technique for Imbalanced Classification Problem
IEEE ACCESS
IF3.6
PSU: Particle Stacking Undersampling Method for Highly Imbalanced Big DataPSU: 高度不平衡大数据的粒子堆叠欠采样方法
IEEE ACCESS
IF3.6

