arrow
返回

Rough set methods in feature selection via submodular function

delete2016-01-30
delete12
PRE
AI
Z
Zhu Xiao-zhong
W
William Zhu *
X
Xinnan Fan
DOI:10.1007/s00500-015-2024-7delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Attribute reduction is an important problem in data mining and machine learning in that it can highlight favorable features and decrease the risk of over-fitting to improve the learning performance. With this regard, rough sets offer interesting opportunities for this problem. Reduct in rough sets is a subspace of attributes/features which are jointly sufficient and individually necessary to satisfy a certain criterion. Excessive attributes may reduce diversity and increase correlation among features, a lower number of attributes may also receive nearly equal to or even higher classification accuracy in some specific classifiers, which motivates us to address dimensionality reduction problems with attribute reduction from the joint viewpoint of the learning performance and the reduct size. In this paper, we propose a new attribute reduction criterion to select lowest attributes while keeping the best performance of the corresponding learning algorithms to some extent. The main contributions of this work are twofold. First, we define the concept of k-approximate-reduct, instead of the limitation to minimum reduct, which provides an important view to reveal the connection between the size of attribute reduct and the learning performance. Second, a greedy algorithm for attribute reduction problems based on mutual information is developed, and submodular functions are used to analyze its convergence. By the property of diminishing return of the submodularity, there is a solid guarantee for the reasonability of the k-approximate-reduct. It is noted that rough sets serve as an effective tool to evaluate both the marginal and joint probability distributions among attributes in mutual information. Extensive experiments in six real-world public datasets from machine learning repository demonstrate that the selected subset by mutual information reduct comes with higher accuracy with less number of attributes when developing classifiers naive Bayes and radial basis function network.
Keyword:
Attribute reduction
Granular computing
Mutual information
Rough set
Submodular function

期刊

Soft Computing 封面图
Soft Computing
IF:
2.5
论文数:
1.0W
被引数:
2.1W

机构

M
Minnan Normal University
学者数:
2.1K
论文数: 1.3K
被引数: 0
H
Hohai University
学者数:
2.3W
论文数: 1.8W
被引数: 2.1W
引用论文

引用论文

Subspace learning for unsupervised feature selection via matrix factorization
err2015-01-01
err142
PREAI
errWang, Shiping; Pedrycz, Witold; Zhu, Qingxin; Zhu, William
err分享
err收藏
Starting a review
err2019-09-20
err0
PREAI
errToby J Lasserson; James Thomas; Julian PT Higgins
err分享
err收藏
err分享
err收藏
Activity of R(+) limonene against Anisakis larvae
err2015-12-01
err0
errOAAI
errFilippo Giarratana; Daniele Muscolino; Felice Panebianco; Andrea Patania; Chiara Benianti; Graziella Ziino; Alessandro Giuffrida
err分享
err收藏
Porous photocatalysts for advanced water purifications用于高级水净化的多孔光催化剂
err2010-01-01
err0
PREAI
errJia Hong Pan; Haiqing Dou; Zhigang Xiong; Chen Xu; Jizhen Ma; X. S. Zhao
err分享
err收藏
err分享
err收藏
err分享
err收藏
Quantized load distribution for tree and bus-connected processors
err2004-07-01
err0
PREAI
errGerassimos Barlas; Bharadwaj Veeravalli
err分享
err收藏
学者 查看更多内容