arrow
返回

Policy Evaluation in Continuous MDPs With Efficient Kernelized Gradient Temporal Difference

delete2021-04-01
delete8
delete
OA
AI
A
Alec Koppel *
G
Garrett Warnell
E
Ethan Stump
P
Peter Stone
A
Alejandro Ribeiro
DOI:10.1109/TAC.2020.3029315delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
We consider policy evaluation in infinite-horizon discounted Markov decision problems with continuous compact state and action spaces. We reformulate this task as a compositional stochastic program with a function-valued decision variable that belongs to a reproducing kernel Hilbert space (RKHS). We approach this problem via a new functional generalization of stochastic quasi-gradient methods operating in tandem with stochastic sparse subspace projections. The result is an extension of gradient temporal difference learning that yields nonlinearly parameterized value function estimates of the solution to the Bellman evaluation equation. We call this method parsimonious kernel gradient temporal difference learning. Our main contribution is a memory-efficient nonparametric stochastic method guaranteed to converge exactly to the Bellman fixed point with probability 1 with attenuating step-sizes under the hypothesis that it belongs to the RKHS. Further, with constant step-sizes and compression budget, we establish mean convergence to a neighborhood and that the value function estimates have finite complexity. In the Mountain Car domain, we observe faster convergence to lower Bellman error solutions than existing approaches with a fraction of the required memory.
Keyword:
Kernel
Complexity theory
Markov processes
Hilbert space
Convergence
Memory management
Automobiles
Iterative learning control
markov processes
optimization methods
stochastic systems
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Automatic Control 封面图
IEEE Transactions on Automatic Control
IF:
7
论文数:
1.3W
被引数:
6.7W

机构

U
us army research, development & engineering command (rdecom)
学者数:
1.6K
论文数: 1.2K
被引数: 0
United States Army 封面图
United States Army
学者数:
5.9K
论文数: 4.3K
被引数: 1.8K
United States Department of Defense 封面图
United States Department of Defense
学者数:
2.8W
论文数: 2.3W
被引数: 172
学者 查看更多机构
引用论文

引用论文

Haloperidol Decanoate Pharmacokinetics in Red Blood Cells and Plasma
err1992-04-01
err0
PREAI
errMAURICE W. DYSKEN; SUCK WON KIM; GOVIND VATASSERY; SUSAN B. JOHNSON; STACY SKARE; LORI HOLDEN; LEE THOMSYCK
err分享
err收藏
The biological transformation of the manufacturing industry – envisioning biointelligent value adding
err2018-01-01
err0
errOAAI
errRobert Miehe; Thomas Bauernhansl; Oliver Schwarz; Andrea Traube; Anselm Lorenzoni; Lara Waltersmann; Johannes Full; Jessica Horbelt; Alexander Sauer
err分享
err收藏
Opening Platforms: How, When and Why?
err2008-01-01
err0
PREAI
errThomas R. Eisenmann; Geoffrey Parker; Marshall W. Van Alstyne
err分享
err收藏
On the Dynamic Performance of Flax Fiber Composite Beams Manufactured at Different Relative Humidity Levels
err2018-08-23
err0
errOAAI
errHuaizhong Li; Abdul Moudood; Wayne Hall; Gaston Francucci; Andreas Öchsner
err分享
err收藏
err分享
err收藏
Use of Amazonian Forest Fragments by Understory Insectivorous Birds
err1995-12-01
err0
PREAI
errPhilip C. Stouffer; Richard O. Bierregaard
err分享
err收藏
Explicit and implicit reinforcement learning across the psychosis spectrum.
err2017-07-01
err0
errOAAI
errDeanna M. Barch; Cameron S. Carter; James M. Gold; Sheri L. Johnson; Ann M. Kring; Angus W. MacDonald; Diego A. Pizzagalli; J. Daniel Ragland; Steven M. Silverstein; Milton E. Strauss
err分享
err收藏
Aminopyrazine CB1 receptor inverse agonists
err2008-06-01
err0
PREAI
errDavid J. Wustrow; George D. Maynard; Jun Yuan; He Zhao; Jianmin Mao; Qin Guo; Mark Kershaw; Jack Hammer; Robbin M. Brodbeck; Kristen E. Near; Dan Zhou; David S. Beers; Bertrand L. Chenard; James E. Krause; Alan J. Hutchison
err分享
err收藏
Surveillance of resected non-small cell lung cancer
err2012-07-21
err0
PREAI
errA. López-González; P. Ibeas Millán; B. Cantos; M. Provencio
err分享
err收藏
学者 查看更多内容