返回
CFMf topic-model: comparison with LDA and Top2Vec
DOI:10.1007/s11192-024-05017-z.png)
摘要
En 中文
Mining the content of scientific publications is increasingly used to investigate the practice of science and the evolution of research domains. Topic models, among which LDA (statistical bag-of-words approach) and Top2Vec (embeddings approach), have notably been shown to provide rich insights into the thematic content of disciplinary fields, their structure and evolution through time. However, improving topic modeling methods remains a major concern. Here we propose an alternative topic-modeling approach based on neural clustering and feature maximization with F1-measure (in short: CFMf). We compare the performance of this approach to LDA and Top2Vec by applying the methods to a reference corpus of full-text philosophy of science articles (N = 16,917). The results reveal significant improvements in terms of coherence measures, independently of the number of topics. Qualitative comparisons show an overall consistency in terms of topical coverage across all three methods, yet with differences: in particular, CFMf appears affected by the presence of a large class while Top2Vec generates some sets of top-words highly difficult to interpret. We discuss these results and highlight upcoming research work.
Keyword:
Topic model
CFMf
LDA
Top2Vec
Clustering
Topic coherence
Topic diversity
期刊
IF:
3.5
论文数:
8.1K
被引数:
2.2W
机构
引用论文
Database of NIH grants using machine-learned categories and graphical clustering
NATURE METHODS
IF32.1
The A-type granitoids: A review of their occurrence and chemical characteristics and speculations on their petrogenesis
Lithos
IF0
An overview of the history of Science of Science in China based on the use of bibliographic and citation data: a new method of analysis based on clustering with feature maximization and contrast graphs
SCIENTOMETRICS
IF3.5
没有更多内容

