arrow
返回

Feature selection and its combination with data over-sampling for multi-class imbalanced datasets

delete2024-03-01
delete5
PRE
AI
C
Chih‐Fong Tsai
K
Kuan-Chen Chen
W
Wei‐Chao Lin *
DOI:10.1016/j.asoc.2024.111267delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Feature selection aims at filtering out some unrepresentative features from a given dataset in order to construct more effective learning models. Furthermore, ensemble feature selection by combining multiple feature selection methods has shown its outperformance over single feature selection. However, the performances of different (ensemble) feature selection methods have not been fully examined over multi -class imbalanced datasets. On the other hand, for class imbalanced datasets, one widely considered solution is to re -balance the datasets by data over -sampling, which generates some synthetic examples for the minority classes. However, the effect of performing (ensemble) feature selection on over -sampling multi -class imbalanced datasets has not been investigated. Therefore, the first research objective is to examine the performances of single and ensemble feature selection methods by fifteen well-known filter, wrapper, and embedded algorithms in terms of classification accuracy. For the second research objective, two orders of combining the feature selection and over -sampling steps are compared in order to find out the best combination procedure as well as the best combined algorithms. The experimental results based on ten different domain datasets containing low to very high feature dimensions show that ensemble feature selection methods slightly perform better than single ones. However, their performance differences are not big. To combine with the Synthetic Minority Oversampling Technique (SMOTE) over -sampling algorithm, performing feature selection first and over -sampling second outperforms the other procedure. Although the best combined algorithms are based on ensemble feature selection, eXtreme Gradient Boosting (XGBoost), as the single best feature selection algorithm, combined with SMOTE provides very similar classification performance to the best combined algorithms. To consider the issues of classification performance and compactional cost, the optimal solution is based on the combined XGBoost and SMOTE.
Keyword:
Feature selection
Ensemble feature selection
Machine learning
Class imbalance learning
Over-sampling

期刊

Applied Soft Computing 封面图
Applied Soft Computing
IF:
6.6
论文数:
1.4W
被引数:
4.8W

机构

N
National Central University
学者数:
1.0W
论文数: 8.6K
被引数: 6.4K
引用论文

引用论文

W′ expenditure and reconstitution during severe intensity constant power exercise: mechanistic insight into the determinants of W′
err2016-09-28
err0
errOAAI
errRyan M. Broxterman; Phillip F. Skiba; Jesse C. Craig; Samuel L. Wilcox; Carl J. Ade; Thomas J. Barstow
err分享
err收藏
err分享
err收藏
Minority oversampling for imbalanced ordinal regression不平衡有序回归的少数过采样
err2019-02-01
err37
PREAI
errZhu, Tuanfei; Lin, Yaping; Liu, Yonghe; Zhang, Wei; Zhang, Jianming
err分享
err收藏
Loss of Atg7 in Endothelial Cells Enhanced Cutaneous Wound Healing in a Mouse Model
err2020-05-01
err0
PREAI
errKe-Cheng Li; Chun-Hui Wang; Jing-Jiang Zou; Chen Qu; Xing-Li Wang; Xing-Song Tian; Hong-Wei Liu; Taixing Cui
err分享
err收藏
Challenges in Automotive Fuel Cells Recycling
err2016-12-01
err0
errOAAI
errRikka Wittstock; Alexandra Pehlken; Michael Wark
err分享
err收藏
err分享
err收藏
学者 查看更多内容