arrow
返回

Text classification method based on self-training and LDA topic models

delete2017-09-01
delete151
PRE
AI
V
Vili Podgorelec
DOI:10.1016/j.eswa.2017.03.020delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Supervised text classification methods are efficient when they can learn with reasonably sized labeled sets. On the other hand, when only a small set of labeled documents is available, semi-supervised methods become more appropriate. These methods are based on comparing distributions between labeled and unlabeled instances, therefore it is important to focus on the representation and its discrimination abilities. In this paper we present the ST LDA method for text classification in a semi-supervised manner with representations based on topic models. The proposed method comprises a semi-supervised text classification algorithm based on self-training and a model, which determines parameter settings for any new document collection. Self-training is used to enlarge the small initial labeled set with the help of information from unlabeled data, We investigate how topic-based representation affects prediction accuracy by performing NBMN and SVM classification algorithms on an enlarged labeled set and then compare the results with the same method on a typical TF-IDF representation. We also compare ST LDA with supervised classification methods and other well-known semi-supervised methods. Experiments were conducted on 11 very small initial labeled sets sampled from six publicly available document collections. The results show that our ST LDA method, when used in combination with NBMN, performed significantly better in terms of classification accuracy than other comparable methods and variations. In this manner, the ST LDA method proved to be a competitive classification method for different text collections when only a small set of labeled instances is available. As such, the proposed ST LDA method may well help to improve text classification tasks, which are essential in many advanced expert and intelligent systems, especially in the case of a scarcity of labeled texts. (C) 2017 Elsevier Ltd. All rights reserved.
Keyword:
Classification
Topic modeling
LDA
Semi-supervised learning
Self-training
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
3.0W
被引数:
10.2W

机构

U
university of maribor
学者数:
4.5K
论文数: 4.1K
被引数: 1
引用论文

引用论文

Statistical topic models for multi-label document classification
err2011-12-29
err229
errOAAI
errRubin, Timothy N.; Chambers, America; Smyth, Padhraic; Steyvers, Mark
err分享
err收藏
err分享
err收藏
Determining the significance and relative importance of parameters of a simulated quenching algorithm using statistical tools
err2011-11-04
err7
PREAI
errCastillo, P. A.; Arenas, M. G.; Rico, N.; Mora, A. M.; Garcia-Sanchez, P.; Laredo, J. L. J.; Merelo, J. J.
err分享
err收藏
err分享
err收藏
err分享
err收藏
A density-based method for adaptive LDA model selection一种基于密度的自适应LDA模型选择方法
err2009-03-01
err552
PREAI
errCao, Juan; Xia, Tian; Li, Jintao; Zhang, Yongdong; Tang, Sheng
err分享
err收藏
学者 查看更多内容