arrow
返回

When Does Cotraining Work in Real Data?

delete2011-05-01
delete69
PRE
AI
J
Jun Du *
C
Charles X. Ling
Z
Zhi‐Hua Zhou
DOI:10.1109/TKDE.2010.158delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Cotraining, a paradigm of semisupervised learning, is promised to alleviate effectively the shortage of labeled examples in supervised learning. The standard two-view cotraining requires the data set to be described by two views of features, and previous studies have shown that cotraining works well if the two views satisfy the sufficiency and independence assumptions. In practice, however, these two assumptions are often not known or ensured (even when the two views are given). More commonly, most supervised data sets are described by one set of attributes (one view). Thus, they need be split into two views in order to apply the standard two-view cotraining. In this paper, we first propose a novel approach to empirically verify the two assumptions of cotraining given two views. Then, we design several methods to split single view data sets into two views, in order to make cotraining work reliably well. Our empirical results show that, given a whole or a large labeled training set, our view verification and splitting methods are quite effective. Unfortunately, cotraining is called for precisely when the labeled training set is small. However, given small labeled training sets, we show that the two cotraining assumptions are difficult to verify, and view splitting is unreliable. Our conclusions for cotraining's effectiveness are mixed. If two views are given, and known to satisfy the two assumptions, cotraining works well. Otherwise, based on small labeled training sets, verifying the assumptions or splitting single view into two views are unreliable; thus, it is uncertain whether the standard cotraining would work or not.
Keyword:
Semisupervised learning
cotraining
sufficiency assumption
independence assumption
view splitting
single-view
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Knowledge and Data Engineering 封面图
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
论文数:
6.8K
被引数:
3.2W

机构

W
western university (university of western ontario)
学者数:
2.9W
论文数: 2.7W
被引数: 33
N
nanjing university
学者数:
7.8W
论文数: 5.6W
被引数: 87
引用论文

引用论文

Estimating multivariate similarity between neuroimaging datasets with sparse canonical correlation analysis: an application to perfusion imaging
err2015-10-13
err0
errOAAI
errMaria J. Rosa; Mitul A. Mehta; Emilio M. Pich; Celine Risterucci; Fernando Zelaya; Antje A. T. S. Reinders; Steve C. R. Williams; Paola Dazzan; Orla M. Doyle; Andre F. Marquand
err分享
err收藏
ENSO signals in East African rainfall seasons
err2000-01-01
err0
PREAI
errMatayo Indeje; Fredrick H.M. Semazzi; Laban J. Ogallo
err分享
err收藏
PP, PYY, and NPY
err1993-01-01
err0
PREAI
errF. Sundler; G. Böttcher; E. Ekblad; R. Håkanson
err分享
err收藏
Text classification from labeled and unlabeled documents using EM
err2000-01-01
err1.9K
errOAAI
errNigam, K; McCallum, AK; Thrun, S; Mitchell, T
err分享
err收藏
err分享
err收藏
学者 查看更多内容