arrow
返回

A multiple association-based unsupervised feature selection algorithm for mixed data sets

delete2023-02-01
delete6
PRE
AI
A
Ayman Taha
A
Ali S. Hadi
S
Susan McKeever *
DOI:10.1016/j.eswa.2022.118718delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Companies have an increasing access to very large datasets within their domain. Analysing these datasets often requires the application of feature selection techniques in order to reduce the dimensionality of the data and prioritize features for downstream knowledge generation tasks. Effective feature selection is a key part of clustering, regression and classification. It presents a myriad of opportunities to improve the machine learning pipeline: eliminating redundant and irrelevant features, reducing model over-fitting, faster model training times and more explainable models. By contrast, and despite the widespread availability and use of categorical data in practice, feature selection for categorical and/or mixed data has received relatively little attention in comparison to numerical data. Furthermore, existing feature selection methods for mixed data are sensitive to number of objects by having nonlinear time complexities with respect to number of objects. In this work, we propose a generic multiple association measure for mixed datasets and a novel feature selection algorithm that uses multiple association across features. Our algorithm is based upon the belief that the most representative chosen set of features should be as diverse and minimally dependent on each other as possible. The proposed algorithm formulates the problem of feature selection as an optimization problem, searching for the set of features that have minimum association amongst them. We present a generic multiple association measure and two associated feature selection algorithms: Naive and Greedy Feature Selection Algorithms called NFSA and GFSA, respectively. Our proposed GFSA algorithm is evaluated on 15 benchmark datasets, and compared to four existing state of the art feature selection techniques. We demonstrate that our approach provides comparable downstream classification performance outperforming other leading techniques on several datasets. Both time complexity analysis and experimental results show that our proposed algorithm significantly reduces the processing time required for unsupervised feature selection algorithms especially for long datasets which have a huge number of objects, whilst also yielding comparable clustering and classification performance. On the other hand, we do not recommend our approach for wide datasets where the number of features is huge with respect to the number of objects e.g., image, text and genome datasets.
Keyword:
Feature selection
Measures of association
Multiple association
Categorical data
Mixed data
Feature engineering

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
3.0W
被引数:
10.2W

机构

A
American University Cairo
学者数:
1.0K
论文数: 803
被引数: 16
E
egyptian knowledge bank (ekb)
学者数:
11.6W
论文数: 9.3W
被引数: 84
C
Cairo University
学者数:
1.4W
论文数: 1.1W
被引数: 1.7W
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
err分享
err收藏
Potential Negative Consequences of Adding Phosphorus‐Based Fertilizers to Immobilize Lead in Soil
err2008-09-01
err0
PREAI
errDouglas W. Kilgour; Rebecca B. Moseley; Mark O. Barnett; Kaye S. Savage; Philip M. Jardine
err分享
err收藏
Security Performance Analysis of Relay Networks Based on κ - μ Shadowed Channels with RHIs and CEEs
err2022-04-13
err0
errOAAI
errJiangfeng Sun; Xiaohong Wang; Yiwei Fang; Xinji Tian; Mingfu Zhu; Jiangtao Ou; Chengyuan Fan
err分享
err收藏
Self-monitoring in the cerebral cortex: Neural responses to small pitch shifts in auditory feedback during speech production
err2018-10-01
err0
errOAAI
errMatthias K. Franken; Frank Eisner; Daniel J. Acheson; James M. McQueen; Peter Hagoort; Jan-Mathijs Schoffelen
err分享
err收藏
学者 查看更多内容