arrow
返回

Variable selection for model-based clustering

delete2006-03-01
delete389
delete
OA
AI
A
Adrian E. Raftery
N
Nema Dean
DOI:10.1198/016214506000000113delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
We consider the problem of variable or feature selection for model-based clustering. The problem of comparing two nested subsets of variables is recast as a model comparison problem and addressed using approximate Bayes factors. A greedy search algorithm is proposed for finding a local optimum in model space. The resulting method selects variables (or features), the number of clusters, and the clustering model simultaneously. We applied the method to several simulated and real examples and found that removing irrelevant variables often improved performance. Compared with methods based on all of the variables, our variable selection method consistently yielded more accurate estimates of the number of groups and lower classification error rates, as well as more parsimonious clustering models and easier visualization of results.
Keyword:
Bayes factor
BIC
feature selection
model-based clustering
unsupervised learning
variable selection

期刊

J
Journal of the American Statistical Association
IF:
3
论文数:
5.2K
被引数:
4.8W

机构

暂无机构信息
引用论文

引用论文

Raman scattering in hydrogenated amorphous silicon under high pressure
err1982-04-01
err0
PREAI
errTakeo Ishidate; Kuon Inoue; Kazuhiko Tsuji; Shigeru Minomura
err分享
err收藏