arrow
返回

Text classification from labeled and unlabeled documents using EM

delete2000-01-01
delete1.9K
delete
OA
AI
K
Kamal Nigam
A
Andrew Kachites McCallum
S
Sebastian Thrun
T
Tom M. Mitchell
DOI:10.1023/A:1007692713085delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
This paper shows that the accuracy of learned text classifiers can be improved by augmenting a small number of labeled training documents with a large pool of unlabeled documents. This is important because in many text classification problems obtaining training labels is expensive, while large quantities of unlabeled documents are readily available. We introduce an algorithm for learning from labeled and unlabeled documents based on the combination of Expectation-Maximization (EM) and a naive Bayes classifier. The algorithm first trains a classifier using the available labeled documents, and probabilistically labels the unlabeled documents. It then trains a new classifier using the labels for all the documents, and iterates to convergence. This basic EM procedure works well when the data conform to the generative assumptions of the model. However these assumptions are often violated in practice, and poor performance can result. We present two extensions to the algorithm that improve classification accuracy under these conditions: (1) a weighting factor to modulate the contribution of the unlabeled data, and (2) the use of multiple mixture components per class. Experimental results, obtained using text from three different real-world tasks, show that the use of unlabeled data reduces classification error by up to 30%.
Keyword:
text classification
Expectation-Maximization
integrating supervised and unsupervised learning
combining labeled and unlabeled data
Bayesian learning
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Machine Learning 封面图
Machine Learning
IF:
2.9
论文数:
2.7K
被引数:
3.4W

机构

暂无机构信息
引用论文

引用论文

暂无论文信息