arrow
Return

Learning noisy linear classifiers via adaptive and selective sampling

delete2010-05-20
delete21
delete
OA
AI
G
Giovanni Cavallanti *
N
Nicolò Cesa‐Bianchi
C
Claudio Gentile
DOI:10.1007/s10994-010-5191-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We introduce efficient margin-based algorithms for selective sampling and filtering in binary classification tasks. Experiments on real-world textual data reveal that our algorithms perform significantly better than popular and similarly efficient competitors. Using the so-called Mammen-Tsybakov low noise condition to parametrize the instance distribution, and assuming linear label noise, we show bounds on the convergence rate to the Bayes risk of a weaker adaptive variant of our selective sampler. Our analysis reveals that, excluding logarithmic factors, the average risk of this adaptive sampler converges to the Bayes risk at rate N (-(1+alpha)(2+alpha)/2(3+alpha)) where N denotes the number of queried labels, and alpha > 0 is the exponent in the low noise condition. For all this convergence rate is asymptotically faster than the rate N (-(1+alpha)/(2+alpha)) achieved by the fully supervised version of the base selective sampler, which queries all labels. Moreover, for alpha -> a (hard margin condition) the gap between the semi- and fully-supervised rates becomes exponential.
Keywords:
Active learning
Selective sampling
Adaptive sampling
Linear classification
Low noise

Journal

Machine Learning cover
Machine Learning
IF:
2.9
Papers:
2.6K
Citations:
3.4W

Organization

U
University of Insubria
Scholars:
6.8K
Papers: 6.0K
Citations: 6.5K
U
University of Milan
Scholars:
5.1W
Papers: 3.9W
Citations: 5.0W