Return
Semantic-Based Evaluation Framework for Topic Models: Integrated Deep Learning and LLM Validation
DOI:10.3745/JIPS.04.0365.png)
Abstract
En 中文
Topic modeling has evolved from statistical methods such as latent Dirichlet allocation (LDA) to neural hybrid models including BERTopic, which utilize bidirectional encoder representations from transformers (BERT) embeddings. However, traditional statistical evaluation metrics overlook the semantic richness of these neural representations, limiting model assessment capabilities. This paper introduces semantic-based evaluation metrics that leverage deep learning embeddings and validates them through both statistical comparison and large language model (LLM)-based assessment. This study evaluated three synthetic datasets with systematically varying topic overlap and one public dataset (20 Newsgroups). Analysis across 9,608 synthetic documents with 45 topics and a stratified sample of 1,000 documents from 20 Newsgroups shows that semantic metrics achieve improved discrimination compared to statistical baselines. Specifically, semantic coherence shows a 38.1% discriminative range versus 5.0% for statistical measures, representing a 7.62 & times; improvement. Semantic distinctiveness achieves 1.57 & times; higher discrimination than statistical methods. Semantic methods also maintain consistent discrimination quality for diversity metrics, with stable progression across similarity levels. LLM assessments, serving as proxies for human judgment, demonstrate inter-model agreement through a weighted three-model ensemble (mean pairwise Spearman rho=0.937) and positive correlation with semantic metrics on public datasets (rho=0.632-0.671). Domain-specific validation and multilingual extension constitute future work.
Keywords:
BERT Embeddings
Contemporary Topic Models
Deep Learning
LLM-based Evaluation
Semantic Evaluation
Metrics
Journal
J
IF:
0.5
Papers:
33
Citations:
0

