返回
Topic representation: Finding more representative words in topic models
DOI:10.1016/j.patrec.2019.01.018.png)
摘要
En 中文
The top word list, i.e., the top-M words with highest marginal probabilities in a given topic, is the standard topic representation in topic models. Most of recent automatical topic labeling algorithms and popular topic quality metrics are based on it. However, we find, empirically, words in this type of top word list are not always representative. The objective of this paper is to find more representative top word lists for topics. To achieve this, we rerank the words in a given topic by further considering marginal probabilities on words over every other topic. The reranking list of top-M words is used to be a novel topic representation for topic models. We investigate three reranking methodologies, using (1) standard deviation weight, (2) standard deviation weight with topic size and (3) Chi Square chi(2) statistic selection. Experimental results on real-world collections indicate that our representations can extract more representative words for topics, agreeing with human judgements. (C) 2019 Elsevier B.V. All rights reserved.
Keyword:
Topic modeling
Topic representation
Topical word representation
Reranking methodology
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.3
论文数:
8.0K
被引数:
1.6W
机构
引用论文
Deep adaptive feature embedding with local sample distributions for person re-identification
PATTERN RECOGNITION
IF7.6
Robust Subspace Clustering for Multi-View Data by Exploiting Correlation Consensus基于相关性一致性的多视图数据鲁棒子空间聚类
没有更多内容

