arrow
返回

A Hybrid Sampling Approach for Imbalanced Binary and Multi-Class Data Using Clustering Analysis

delete2022-01-01
delete12
delete
OA
AI
A
Abdul Sattar Palli *
J
Jafreezal Jaafar
M
Manzoor Ahmed Hashmani
H
Heitor Murilo Gomes
A
Abdul Rehman Gilal
DOI:10.1109/ACCESS.2022.3218463delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Unequal data distribution among different classes usually cause a class imbalance problem. Due to the class imbalance, the classification models become biased toward the majority class and misclassify the minority class. Class imbalance issue becomes more complex when it occurs in multi-class data. The most common method to handle the class imbalance is data resampling that involves either over-sampling minority class instances or under-sampling majority class instances. In the case of under-sampling, there is a chance of losing some crucial information, whereas over-sampling can cause an overfitting problem. Therefore, we propose a novel Cluster-based Hybrid Sampling for Imbalance Data (CBHSID) approach to address these issues. The CBHSID calculates the mean of the data observations based on the number of classes. It uses the calculated mean as a threshold value to segregate majority and minority classes. CBHSID applies affinity propagation cluster analysis to each class to create sub-clusters and calculates the distance of each data item of sub-cluster using centroid mean. CBHSID removes data observations that are away from the center of sub-cluster during under-sampling. On the other hand, during the over-sampling, it generates synthetic samples using data observations near to the center of sub-cluster. We compared CBHSID with a few state-of-the-art data balancing methods on 12 binary and 4 multi-class benchmark datasets. Based on Geometric-Mean (G-Mean), Recall, and F1-score, our method outperformed the other compared methods on 14 datasets out of 16. Results also revealed that CBHSID is suitable for addressing class imbalance issues in both binary and multi-class classifications. In the current state, we have only validated CBHSID on stationary data streams. Consequently, CBHSID can further be tested on non-stationary data streams in online learning environments.
Keyword:
Class imbalance
classification
clustering analysis
binary class
multi-class

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

U
Universiti Teknologi Petronas
学者数:
5.4K
论文数: 4.6K
被引数: 5.9K
引用论文

引用论文

Isolation and characterization of drought responsiveEcDehydrin7gene from finger millet (Eleusine coracana(L.) Gaertn.)
err2014-01-01
err0
PREAI
errRajiv Kumar Singh; Mullapudi Lakshmi Venkata Phanindra; Vivek Kumar Singh; Sonam; Sanagala Raghavendrarao; Amolkumar U. Solanke; Polumetla Ananda Kumar
err分享
err收藏
err分享
err收藏
A comprehensive data level analysis for cancer diagnosis on imbalanced data
err2019-02-01
err194
PREAI
errFotouhi, Sara; Asadi, Shahrokh; Kattan, Michael W.
err分享
err收藏
Saline with benzyl alcohol as intradermal anesthesia for intravenous line placement in children
err1998-04-01
err0
PREAI
errJOEL A. FEIN; CHRIS R. BOARDMAN; SUE STEVENSON; STEVEN M. SELBST
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
学者 查看更多内容