arrow
Return

Optimal Sampling Rate Selection for Parallel Hybrid Sampling Framework of Imbalanced Data Classification

delete2025-12-25
delete0
PRE
AI
Z
Zhuo Zhao
M
Ming Zheng *
M
Ma, Fanhao
Y
Yi Zheng
DOI:10.1002/cpe.70462delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Imbalanced data classification is one of the challenges in the field of data mining and machine learning. At present, the main method to solve imbalanced data classification issues from the data level is resampling. Hybrid sampling is widely used because it can avoid the problem of overfitting or mistakenly deleting the most useful samples when using oversampling or undersampling alone. However, the current hybrid sampling methods are mostly implemented in serial, which has the problems of excessive time cost and mutual influence. Meanwhile, few studies consider the automatic determination of the oversampling rate and undersampling rate in hybrid sampling methods but use the default sampling rate. Therefore, this study proposes a novel optimal sampling rate selection for a parallel hybrid sampling framework of imbalanced data classification. At the same time, we improved the differential evolution algorithm to optimize the sampling rate for the parallel hybrid sampling framework to obtain more stable classification performance. Experiments show that the SRDE specifically designed for the parallel hybrid sampling framework is superior to other heuristic algorithms on imbalanced data classification, and the results of statistical test experiments also verified this result.
Keywords:
data level
heuristic algorithm
hybrid sampling
imbalanced data

Journal

C
CONCURRENCY AND COMPUTATION-PRACTICE & EXPERIENCE
IF:
1.5
Papers:
473
Citations:
0

Organization

A
Anhui Normal University
Scholars:
7.0K
Papers: 4.6K
Citations: 6.8K