arrow
返回

Learning aspect models with partially labeled data

delete2011-01-01
delete3
delete
OA
AI
A
Anastasia Krithara *
C
Cyril Goutte
J
Jean-Michel Renders
DOI:10.1016/j.patrec.2010.09.004delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
In this paper we address the problem of learning aspect models with partially labeled data for the task of document categorization The motivation of this work is to take advantage of the amount of available unlabeled data together with the set of labeled examples to learn latent models whose structure and underlying hypotheses take more accurately Into account the document generation process compared to other mixture-based generative models We present one semi-supervised variant of the Probabilistic Latent Semantic Analysis (PLSA) model (Hofmann 2001) In our approach we try to capture the possible data mislabeling errors which occur during the training of our model This is done by iteratively assigning class labels to unlabeled examples using the current aspect model and re-estimating the probabilities of the mislabeling errors We perform experiments over the 20Newsgroups WebKB and Reuters document collections as well as over a real world dataset coming from a Business Group of Xerox and show the effectiveness of our approach compared to a semi-supervised version of Naive Bayes another semi-supervised version of PLSA and to transductive Support Vector Machines (C) 2010 Elsevier B V All rights reserved
Keyword:
Semi supervised learning
Aspect models
Document categorization
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Pattern Recognition Letters 封面图
Pattern Recognition Letters
IF:
3.3
论文数:
7.9K
被引数:
1.6W

机构

X
xerox
学者数:
226
论文数: 183
被引数: 0
N
National Research Council Canada
学者数:
7.9K
论文数: 7.9K
被引数: 6.8K
S
Sorbonne Universite
学者数:
6.2W
论文数: 4.5W
被引数: 605
学者 查看更多机构
引用论文

引用论文

Text classification from labeled and unlabeled documents using EM
err2000-01-01
err1.9K
errOAAI
errNigam, K; McCallum, AK; Thrun, S; Mitchell, T
err分享
err收藏
学者 查看更多内容