返回
Clustering scientific documents with topic modeling
DOI:10.1007/s11192-014-1321-8.png)
摘要
En 中文
Topic modeling is a type of statistical model for discovering the latent topics'' that occur in a collection of documents through machine learning. Currently, latent Dirichlet allocation (LDA) is a popular and common modeling approach. In this paper, we investigate methods, including LDA and its extensions, for separating a set of scientific publications into several clusters. To evaluate the results, we generate a collection of documents that contain academic papers from several different fields and see whether papers in the same field will be clustered together. We explore potential scientometric applications of such text analysis capabilities.
Keyword:
Topic modeling
Text analysis
Atent dirichlet allocation
期刊
IF:
3.5
论文数:
8.1K
被引数:
2.2W
机构
引用论文
Overlaying communities and topics: an analysis on publication networks重叠的社区和主题: 对出版网络的分析
SCIENTOMETRICS
IF3.5

