返回
Improving document clustering in a learned concept space
DOI:10.1016/j.ipm.2009.09.007.png)
摘要
En 中文
Most document clustering algorithms operate in a high dimensional bag-of-words space. The inherent presence of noise in such representation obviously degrades the performance of most of these approaches. In this paper we investigate an unsupervised dimensionality reduction technique for document clustering. This technique is based upon the assumption that terms co-occurring in the same context with the same frequencies are semantically related. On the basis of this assumption we first find term clusters using a classification version of the EM algorithm. Documents are then represented in the space of these term clusters and a multinomial mixture model (mm) is used to build document clusters. We empirically show on four document collections, Reuters-21578, Reuters RCV2-French, 20Newsgroups and WebKB, that this new text representation noticeably increases the performance of the mm model. By relating the proposed approach to the Probabilistic Latent Semantic Analysis (PLSA) model we further propose an extension of the latter in which an extra latent variable allows the model to co-cluster documents and terms simultaneously. We show on these four datasets that the proposed extended version of the PLSA model produces statistically significant improvements with respect to two clustering measures over all variants of the original PLSA and the mm models. (C) 2009 Elsevier Ltd. All rights reserved.
Keyword:
Document clustering
Aspect models
Concept learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
I
IF:
6.9
论文数:
5.2K
被引数:
1.4W
机构
引用论文
Prediction of Pregabalin-Mediated Pain Response by Severity of Sleep Disturbance in Patients with Painful Diabetic Neuropathy and Post-Herpetic Neuralgia疼痛性糖尿病神经病变和带状疱疹后神经痛患者睡眠障碍严重程度预测普瑞巴林介导的疼痛反应

