arrow
返回

Unsupervised feature selection method based on iterative similarity graph factorization and clustering by modularity

delete2022-12-01
delete4
PRE
AI
M
Marcos de Oliveira *
S
Sergio Queiroz
F
Francisco de A.T. de Carvalho
DOI:10.1016/j.eswa.2022.118092delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Feature selection is an important research area aimed at eliminating unwanted features from high-dimensional datasets. Feature selection methods are categorized according to the availability of labels. Supervised methods usually select features that have high correlations with the data labels, however, this cannot be done by unsupervised methods, since the labels are not available, making the development of them even more challenging. Some unsupervised methods proposed in the literature get around this problem by generating pseudo-labels, using clustering techniques, and then making feature selection as in the supervised scenario. There are also methods in which the selection is not guided by a clustering criterion, generally performing some kind of reconstruction of the original dataset in low dimensionality. However, in both approaches a local structure is generated, usually starting from a similarity graph built by using the entire set of features. This may compromise the results of the methods when there are many irrelevant and noisy features, which makes it difficult to reveal patterns or an initial representation of the data. Another drawback of some current methods is the large amount of hyper-parameters, or the computational time required for a single execution, which can render such methods unfeasible for some large datasets. In order to address these problems, it is proposed in this work a new unsupervised feature selection method, called KNMFS, which performs the scoring of the features in low computational time using a three-step procedure: (1) a similarity graph learning step based on non-negative iterative matrix factorization together with a Gaussian filter to mitigate the noise effects caused by irrelevant features; (2) a clustering step from the learned graph using a modularity optimization strategy and (3) a random forest algorithm is applied to compute the scores based on the labels generated by the clustering step. To verify the effectiveness of the proposed method, experiments in real and synthetic datasets were conducted. In both cases, the obtained results showed that the KNMFS method, compared to other state-of-the-art methods, obtained good results according to the metrics of Accuracy, ARI and NMI. Friedman's statistical tests were also performed to give stronger evidence to the reported results.
Keyword:
Featureselection
Unsupervisedlearning
Low-rankapproximation
Graphmodularity

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
2.9W
被引数:
10.2W

机构

U
Universidade Federal de Pernambuco
学者数:
1.3W
论文数: 7.3K
被引数: 5.3K
引用论文

引用论文

err分享
err收藏
Subspace clustering guided unsupervised feature selection
err2017-06-01
err187
PREAI
errZhu, Pengfei; Zhu, Wencheng; Hu, Qinghua; Zhang, Changqing; Zuo, Wangmeng
err分享
err收藏
Deep feature selection using a teacher-student network基于师生网络的深度特征选择
err2020-03-01
err41
errOAAI
errMirzaei, Ali; Pourahmadi, Vahid; Soltani, Mehran; Sheikhzadeh, Hamid
err分享
err收藏
Low Rank Regularization: A review
err2021-04-01
err56
errOAAI
errHu, Zhanxuan; Nie, Feiping; Wang, Rong; Li, Xuelong
err分享
err收藏
学者 查看更多内容