返回
TopicStriKer: A topic kernels-powered approach for text classification
DOI:10.1016/j.rineng.2023.100949.png)
摘要
En 中文
Topic models are unsupervised machine learning techniques that output clusters of topics represented as co-occurring words with their associated probability distributions. Topic modeling algorithms find latent themes from large document collections by understanding their context. On the other hand, string kernels are supervised machine-learning techniques that quantify string similarities without explicit string encoding. We propose TopicStriKer, a model combining the advantages of unsupervised topic modeling with supervised string kernels for text classification tasks. The co-occurring topic words per topic and topic proportions per document obtained are used to reduce the document corpus to a topic-word sequence. This reduced representation is then used for text classification with the aid of string kernels, significantly improving accuracy and reducing training time. Experiments on the bag-of-words kernel-based string embeddings using the proposed algorithm outperform the traditional text classification approaches. This work extensively compares string kernels with topic modeling on various performance metrics to establish our findings.
Keyword:
Topic modeling
String kernels
Text classification
String embedding
Topic sequence
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.9
论文数:
1.2W
被引数:
1.7W
机构
引用论文
A new topic modeling based approach for aspect extraction in aspect based sentiment analysis: SS-LDA基于方面的情感分析中一种新的基于主题建模的方面提取方法: ss-lda

