返回
Interpretable Topic Extraction and Word Embedding Learning Using Non-Negative Tensor DEDICOM
DOI:10.3390/make3010007.png)
摘要
En 中文
Unsupervised topic extraction is a vital step in automatically extracting concise contentual information from large text corpora. Existing topic extraction methods lack the capability of linking relations between these topics which would further help text understanding. Therefore we propose utilizing the Decomposition into Directional Components (DEDICOM) algorithm which provides a uniquely interpretable matrix factorization for symmetric and asymmetric square matrices and tensors. We constrain DEDICOM to row-stochasticity and non-negativity in order to factorize pointwise mutual information matrices and tensors of text corpora. We identify latent topic clusters and their relations within the vocabulary and simultaneously learn interpretable word embeddings. Further, we introduce multiple methods based on alternating gradient descent to efficiently train constrained DEDICOM algorithms. We evaluate the qualitative topic modeling and word embedding performance of our proposed methods on several datasets, including a novel New York Times news dataset, and demonstrate how the DEDICOM algorithm provides deeper text analysis than competing matrix factorization approaches.
Keyword:
matrix factorization
tensor factorization
word embeddings
topic modeling
NLP
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
M
IF:
6
论文数:
834
被引数:
1.8K
机构
引用论文
Cobalt and Copper Composite Oxides as Efficient Catalysts for Preferential Oxidation of CO in H2-Rich Stream钴和铜复合氧化物作为H2-Rich流中CO优先氧化的有效催化剂
Oxidative N-heterocyclic carbene catalyzed stereoselective annulation of simple aldehydes and 5-alkenyl thiazolones氧化N-杂环卡宾催化的简单醛和5-烯基噻唑酮的立体选择性环化

