返回
Undersampling based on generalized learning vector quantization and natural nearest neighbors for imbalanced data
DOI:10.1007/s13042-024-02261-w.png)
摘要
En 中文
Imbalanced datasets can adversely affect classifier performance. Conventional undersampling approaches may lead to the loss of essential information, while oversampling techniques could introduce noise. To address this challenge, we propose an undersampling algorithm called GLNDU (Generalized Learning Vector Quantization and Natural Nearest Neighbors-based Undersampling). GLNDU utilizes Generalized Learning Vector Quantization (GLVQ) for computing the centroids of positive and negative instances. It also utilizes the concept of Natural Nearest Neighbors to identify majority-class instances in the overlapping region of the centroids of minority-class instances. Afterwards, these majority-class instances are removed, resulting in a new balanced training dataset that is used to train a foundational classifier. We conduct extensive experiments on 29 publicly available datasets, evaluating the performance using AUC and G_mean values. GLNDU demonstrates significant advantages over established methods such as SVM, CART, and KNN across different types of classifiers. Additionally, the results of the Friedman ranking and Nemenyi post-hoc test provide additional support for the findings obtained from the experiments.
Keyword:
Imbalanced data
GLVQ
Natural nearest neighbors
Resampling techniques
Class overlap
期刊
IF:
2.7
论文数:
3.2K
被引数:
5.6K
机构
引用论文
Whole-body MRI in generalized cystic lymphangiomatosis in the pediatric population: diagnosis, differential diagnoses, and follow-up全身MRI在儿童泛发性囊性淋巴管瘤病中的应用:诊断、鉴别诊断及随访
Weighted Ensemble with one-class Classification and Over-sampling and Instance selection (WECOI): An approach for learning from imbalanced data streams具有单类分类,过采样和实例选择 (WECOI) 的加权集成: 一种从不平衡数据流中学习的方法
Arsenic removal from aqueous solutions by adsorption using novel MIL-53(Fe) as a highly efficient adsorbent使用新型MIL-53(Fe) 作为高效吸附剂通过吸附从水溶液中去除砷
RSC Advances
IF0
A framework based on local cores and synthetic examples generation for self-lab ele d semi-supervise d classification
PATTERN RECOGNITION
IF7.6
A novel progressively undersampling method based on the density peaks sequence for imbalanced data一种基于密度峰值序列的不平衡数据渐进欠采样方法

