arrow
Return

Bayesian optimistic Kullback-Leibler exploration

delete2018-12-19
delete2
delete
OA
AI
K
Kanghoon Lee
G
Geon-Hyeong Kim
P
Pedro A. Ortega
D
Daniel D. Lee
K
Kee-Eung Kim *
DOI:10.1007/s10994-018-5767-4delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
We consider a Bayesian approach to model-based reinforcement learning, where the agent uses a distribution of environment models to find the action that optimally trades off exploration and exploitation. Unfortunately, it is intractable to find the Bayes-optimal solution to the problem except for restricted cases. In this paper, we present BOKLE, a simple algorithm that uses Kullback-Leibler divergence to constrain the set of plausible models for guiding the exploration. We provide a formal analysis that this algorithm is near Bayes-optimal with high probability. We also show an asymptotic relation between the solution pursued by BOKLE and a well-known algorithm called Bayesian exploration bonus. Finally, we show experimental results that clearly demonstrate the exploration efficiency of the algorithm.
Keywords:
Model-based Bayesian reinforcement learning
Bayes-adaptive Markov decision process
PAC-BAMDP
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Machine Learning cover
Machine Learning
IF:
2.9
Papers:
2.6K
Citations:
3.4W

Organization

G
Google Incorporated
Scholars:
3.5K
Papers: 1.8K
Citations: 8