arrow
Return

Variational learning from implicit bandit feedback

delete2021-07-09
delete0
delete
OA
AI
Q
Quoc-Tuan Truong
H
Hady W. Lauw *
DOI:10.1007/s10994-021-06028-0delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Recommendations are prevalent in Web applications (e.g., search ranking, item recommendation, advertisement placement). Learning from bandit feedback is challenging due to the sparsity of feedback limited to system-provided actions. In this work, we focus on batch learning from logs of recommender systems involving both bandit and organic feedbacks. We develop a probabilistic framework with a likelihood function for estimating not only explicit positive observations but also implicit negative observations inferred from the data. Moreover, we introduce a latent variable model for organic-bandit feedbacks to robustly capture user preference distributions. Next, we analyze the behavior of the new likelihood under two scenarios, i.e., with and without counterfactual re-weighting. For speedier item ranking, we further investigate the possibility of using Maximum-a-Posteriori (MAP) estimate instead of Monte Carlo (MC)-based approximation for prediction. Experiments on both real datasets as well as data from a simulation environment show substantial performance improvements over comparable baselines.
Keywords:
Variational learning
Bandit feedback
Recommender systems
Computational advertising
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Machine Learning cover
Machine Learning
IF:
2.9
Papers:
2.6K
Citations:
3.4W

Organization

S
Singapore Management University
Scholars:
1.5K
Papers: 2.5K
Citations: 3.5K