arrow
Return

Model selection in reinforcement learning

delete2011-06-11
delete32
delete
OA
AI
A
Amir‐massoud Farahmand
C
Csaba Szepesvári *
DOI:10.1007/s10994-011-5254-7delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We consider the problem of model selection in the batch (offline, non-interactive) reinforcement learning setting when the goal is to find an action-value function with the smallest Bellman error among a countable set of candidates functions. We propose a complexity regularization-based model selection algorithm, BERMIN, and prove that it enjoys an oracle-like property: the estimator's error differs from that of an oracle, who selects the candidate with the minimum Bellman error, by only a constant factor and a small remainder term that vanishes at a parametric rate as the number of samples increases. As an application, we consider a problem when the true action-value function belongs to an unknown member of a nested sequence of function spaces. We show that under some additional technical conditions BERMIN leads to a procedure whose rate of convergence, up to a constant factor, matches that of an oracle who knows which of the nested function spaces the true action-value function belongs to, i.e., the procedure achieves adaptivity.
Keywords:
Reinforcement learning
Model selection
Complexity regularization
Adaptivity
Offline learning
Off-policy learning
Finite-sample bounds

Journal

Machine Learning cover
Machine Learning
IF:
2.9
Papers:
2.6K
Citations:
3.4W

Organization

U
university of alberta
Scholars:
5.1W
Papers: 4.9W
Citations: 65