arrow
Return

Policy Evaluation in Continuous MDPs With Efficient Kernelized Gradient Temporal Difference

delete2021-04-01
delete8
delete
OA
AI
A
Alec Koppel *
G
Garrett Warnell
E
Ethan Stump
P
Peter Stone
A
Alejandro Ribeiro
DOI:10.1109/TAC.2020.3029315delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
We consider policy evaluation in infinite-horizon discounted Markov decision problems with continuous compact state and action spaces. We reformulate this task as a compositional stochastic program with a function-valued decision variable that belongs to a reproducing kernel Hilbert space (RKHS). We approach this problem via a new functional generalization of stochastic quasi-gradient methods operating in tandem with stochastic sparse subspace projections. The result is an extension of gradient temporal difference learning that yields nonlinearly parameterized value function estimates of the solution to the Bellman evaluation equation. We call this method parsimonious kernel gradient temporal difference learning. Our main contribution is a memory-efficient nonparametric stochastic method guaranteed to converge exactly to the Bellman fixed point with probability 1 with attenuating step-sizes under the hypothesis that it belongs to the RKHS. Further, with constant step-sizes and compression budget, we establish mean convergence to a neighborhood and that the value function estimates have finite complexity. In the Mountain Car domain, we observe faster convergence to lower Bellman error solutions than existing approaches with a fraction of the required memory.
Keywords:
Kernel
Complexity theory
Markov processes
Hilbert space
Convergence
Memory management
Automobiles
Iterative learning control
markov processes
optimization methods
stochastic systems
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Automatic Control cover
IEEE Transactions on Automatic Control
IF:
7
Papers:
1.3W
Citations:
6.7W

Organization

United States Army cover
United States Army
Scholars:
5.9K
Papers: 4.3K
Citations: 1.8K
United States Department of Defense cover
United States Department of Defense
Scholars:
2.8W
Papers: 2.3W
Citations: 172
researcher View more organizations
Cited Papers

Cited Papers

Haloperidol Decanoate Pharmacokinetics in Red Blood Cells and Plasma
err1992-04-01
err0
PREAI
errMAURICE W. DYSKEN; SUCK WON KIM; GOVIND VATASSERY; SUSAN B. JOHNSON; STACY SKARE; LORI HOLDEN; LEE THOMSYCK
errShare
errSave
The biological transformation of the manufacturing industry – envisioning biointelligent value adding
err2018-01-01
err0
errOAAI
errRobert Miehe; Thomas Bauernhansl; Oliver Schwarz; Andrea Traube; Anselm Lorenzoni; Lara Waltersmann; Johannes Full; Jessica Horbelt; Alexander Sauer
errShare
errSave
Opening Platforms: How, When and Why?
err2008-01-01
err0
PREAI
errThomas R. Eisenmann; Geoffrey Parker; Marshall W. Van Alstyne
errShare
errSave
On the Dynamic Performance of Flax Fiber Composite Beams Manufactured at Different Relative Humidity Levels
err2018-08-23
err0
errOAAI
errHuaizhong Li; Abdul Moudood; Wayne Hall; Gaston Francucci; Andreas Öchsner
errShare
errSave
errShare
errSave
Use of Amazonian Forest Fragments by Understory Insectivorous Birds
err1995-12-01
err0
PREAI
errPhilip C. Stouffer; Richard O. Bierregaard
errShare
errSave
Explicit and implicit reinforcement learning across the psychosis spectrum.
err2017-07-01
err0
errOAAI
errDeanna M. Barch; Cameron S. Carter; James M. Gold; Sheri L. Johnson; Ann M. Kring; Angus W. MacDonald; Diego A. Pizzagalli; J. Daniel Ragland; Steven M. Silverstein; Milton E. Strauss
errShare
errSave
Aminopyrazine CB1 receptor inverse agonists
err2008-06-01
err0
PREAI
errDavid J. Wustrow; George D. Maynard; Jun Yuan; He Zhao; Jianmin Mao; Qin Guo; Mark Kershaw; Jack Hammer; Robbin M. Brodbeck; Kristen E. Near; Dan Zhou; David S. Beers; Bertrand L. Chenard; James E. Krause; Alan J. Hutchison
errShare
errSave
Surveillance of resected non-small cell lung cancer
err2012-07-21
err0
PREAI
errA. López-González; P. Ibeas Millán; B. Cantos; M. Provencio
errShare
errSave
researcher View more