arrow
Return

Variable Selection Via Thompson Sampling

delete2021-07-06
delete8
delete
OA
AI
Y
Yi Liu
V
Veronika Ročková *
DOI:10.1080/01621459.2021.1928514delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Thompson sampling is a heuristic algorithm for the multi-armed bandit problem which has a long tradition in machine learning. The algorithm has a Bayesian spirit in the sense that it selects arms based on posterior samples of reward probabilities of each arm. By forging a connection between combinatorial binary bandits and spike-and-slab variable selection, we propose a stochastic optimization approach to subset selection called Thompson variable selection (TVS). TVS is a framework for interpretable machine learning which does not rely on the underlying model to be linear. TVS brings together Bayesian reinforcement and machine learning in order to extend the reach of Bayesian subset selection to nonparametric models and large datasets with very many predictors and/or very many observations. Depending on the choice of a reward, TVS can be deployed in offline as well as online setups with streaming data batches. Tailoring multiplay bandits to variable selection, we provide regret bounds without necessarily assuming that the arm mean rewards be unrelated. We show a very strong empirical performance on both simulated and real data. Unlike deterministic optimization methods for spike-and-slab variable selection, the stochastic nature makes TVS less prone to local convergence and thereby more robust.
Keywords:
BART
Combinatorial bandits
Interpretable machine learning
Spike-and-slab
Thompson sampling
Variable selection
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

J
Journal of the American Statistical Association
IF:
3
Papers:
5.1K
Citations:
4.8W

Organization

U
university of chicago
Scholars:
4.4W
Papers: 3.7W
Citations: 80