Return
BATCHED BANDIT PROBLEMS
DOI:10.1214/15-AOS1381.png)
Abstract
En 中文
Motivated by practical applications, chiefly clinical trials, we study the regret achievable for stochastic bandits under the constraint that the employed policy must split trials into a small number of batches. We propose a simple policy, and show that a very small number of batches gives close to minimax optimal regret bounds. As a byproduct, we derive optimal policies with low switching cost for stochastic bandits.
Keywords:
Multi-armed bandit problems
regret bounds
batches
multi-phase allocation
grouped clinical trials
sample size determination
switching cost
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
3.7
Papers:
2.8K
Citations:
2.9W

