arrow
返回

Multiclass classification with bandit feedback using adaptive regularization

delete2012-10-24
delete36
delete
OA
AI
K
Koby Crammer *
C
Claudio Gentile
DOI:10.1007/s10994-012-5321-8delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
We present a new multiclass algorithm in the bandit framework, where after making a prediction, the learning algorithm receives only partial feedback, i.e., a single bit indicating whether the predicted label is correct or not, rather than the true label. Our algorithm is based on the second-order Perceptron, and uses upper-confidence bounds to trade-off exploration and exploitation, instead of random sampling as performed by most current algorithms. We analyze this algorithm in a partial adversarial setting, where instances are chosen adversarially, while the labels are chosen according to a linear probabilistic model which is also chosen adversarially. We show a regret of , which improves over the current best bounds of in the fully adversarial setting. We evaluate our algorithm on nine real-world text classification problems and on four vowel recognition tasks, often obtaining state-of-the-art results, even compared with non-bandit online algorithms, especially when label noise is introduced.
Keyword:
Online learning
Upper confidence bound
Regret

期刊

Machine Learning 封面图
Machine Learning
IF:
2.9
论文数:
2.6K
被引数:
3.4W

机构

U
University of Insubria
学者数:
6.8K
论文数: 6.0K
被引数: 6.5K
T
Technion Israel Institute of Technology
学者数:
1.6W
论文数: 1.5W
被引数: 2.0W