arrow
Return

A Sampling-Based Gittins Index Approximation

delete2026-03-01
delete0
PRE
AI
B
Baas, Stef *
R
Richard J. Boucherie
B
Braaksma, Aleida
DOI:10.1287/moor.2023.0225delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
A sampling-based method is introduced to approximate the Gittins index for a general family of alternative bandit processes. The approximation consists of a truncation of the optimization horizon and support for the immediate rewards, an optimal stopping value approximation, and a stochastic approximation procedure. Finite-time error bounds are given for the three approximations, leading to a procedure to construct a confidence interval for the Gittins index using a finite number of Monte Carlo samples as well as an epsilon-optimal policy for the family of alternative bandit processes. Proofs are given for almost sure convergence and a central limit theorem for the sampling-based Gittins index approximation. In a numerical study, the quality of the approximation is verified for the Bernoulli bandit and the Gaussian bandit with known variance, and the method is shown to significantly outperform Thompson sampling and the Bayesian upper-confidencebound algorithms for a novel random effects multi-armed bandit.
Keywords:
stochastic approximation
multi-armed bandits
optimal stopping
Bayesian computation
Markov decision processes

Journal

M
Mathematics of Operations Research
IF:
1.9
Papers:
77
Citations:
0

Organization

U
university of twente
Scholars:
1.5W
Papers: 1.4W
Citations: 9