返回
Learning aspect models with partially labeled data
DOI:10.1016/j.patrec.2010.09.004.png)
摘要
En 中文
In this paper we address the problem of learning aspect models with partially labeled data for the task of document categorization The motivation of this work is to take advantage of the amount of available unlabeled data together with the set of labeled examples to learn latent models whose structure and underlying hypotheses take more accurately Into account the document generation process compared to other mixture-based generative models We present one semi-supervised variant of the Probabilistic Latent Semantic Analysis (PLSA) model (Hofmann 2001) In our approach we try to capture the possible data mislabeling errors which occur during the training of our model This is done by iteratively assigning class labels to unlabeled examples using the current aspect model and re-estimating the probabilities of the mislabeling errors We perform experiments over the 20Newsgroups WebKB and Reuters document collections as well as over a real world dataset coming from a Business Group of Xerox and show the effectiveness of our approach compared to a semi-supervised version of Naive Bayes another semi-supervised version of PLSA and to transductive Support Vector Machines (C) 2010 Elsevier B V All rights reserved
Keyword:
Semi supervised learning
Aspect models
Document categorization
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.3
论文数:
7.9K
被引数:
1.6W
机构
引用论文
Neoliberal Reforms in an Emerging Democracy: The Case of the Privatization of Public Enterprises in Nigeria, 1999–2014新自由主义改革在新兴民主国家:尼日利亚国有企业私有化案例研究(1999—2014)
Unsupervised learning by probabilistic latent semantic analysis基于概率潜在语义分析的无监督学习
MACHINE LEARNING
IF2.9

