返回
A novel two-phase clustering-based under-sampling method for imbalanced classification problems
DOI:10.1016/j.eswa.2022.119003.png)
摘要
En 中文
Classification problems with imbalanced data are challenging because traditional classifiers tend to misclassify minority samples. This paper introduces a novel two-phase method in the categories of under-sampling data-level and ensemble-based approaches to tackle these problems. Compared to classic k-means clustering, the main property of our method is that the majority class is partitioned into clusters so that there are no minority samples in the convex-hull of the majority samples of each cluster, and at the same time, the size of each cluster is controlled. Thus, it is expected that by using our method, the general pattern of data in the feature space is kept, and the possibility of changing the data distribution is reduced. Computational results over a variety of imbal-anced classification datasets confirm the superiority of our method over the existing methods from different metrics.
Keyword:
Imbalanced data
Under-sampling
Clustering
Convex-hull
Ensemble learning
期刊
IF:
7.5
论文数:
2.9W
被引数:
10.2W
机构
引用论文
EAMA: Empirically adjusted meta-analysis for large-scale simultaneous hypothesis testing in genomic experiments
PLOS ONE
IF0

