arrow
返回

Importance sampling policy gradient algorithms in reproducing kernel Hilbert space

delete2017-10-10
delete5
PRE
AI
T
Tuyen Pham Le
V
Vien Anh Ngo
P
P. Marlith Jaramillo
T
TaeChoong Chung *
DOI:10.1007/s10462-017-9579-xdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Modeling policies in reproducing kernel Hilbert space (RKHS) offers a very flexible and powerful new family of policy gradient algorithms called RKHS policy gradient algorithms. They are designed to optimize over a space of very high or infinite dimensional policies. As a matter of fact, they are known to suffer from a large variance problem. This critical issue comes from the fact that updating the current policy is based on a functional gradient that does not exploit all old episodes sampled by previous policies. In this paper, we introduce a generalized RKHS policy gradient algorithm that integrates the following important ideas: (i) policy modeling in RKHS; (ii) normalized importance sampling, which helps reduce the estimation variance by reusing previously sampled episodes in a principled way; and (iii) regularization terms, which avoid updating the policy too over-fit to sampled data. In the experiment section, we provide an analysis of the proposed algorithms through bench-marking domains. The experiment results show that the proposed algorithm can still enjoy a powerful policy modeling in RKHS and achieve more data-efficiency.
Keyword:
Reproducing kernel Hilbert space
Policy search
Reinforcement learning
Importance sampling
Policy gradient
Non-parametric
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Artificial Intelligence Review 封面图
Artificial Intelligence Review
IF:
13.9
论文数:
6.1K
被引数:
1.9W

机构

Q
Queen's University Belfast
学者数:
1.6W
论文数: 1.7W
被引数: 2.5W
K
kyung hee university
学者数:
2.3W
论文数: 2.2W
被引数: 234
引用论文

引用论文

Reinforcement learning of motor skills with policy gradients
err2008-05-01
err628
PREAI
errPeters, Jan; Schaal, Stefan
err分享
err收藏
err分享
err收藏
What Is the Relationship of Fear Avoidance to Physical Function and Pain Intensity in Injured Athletes?
err2018-02-16
err0
errOAAI
errStefan F. Fischerauer; Mojtaba Talaei-Khoei; Rens Bexkens; David C. Ring; Luke S. Oh; Ana-Maria Vranceanu
err分享
err收藏
Kernel matching pursuit内核匹配追踪
err2002-01-01
err242
errOAAI
errVincent, P; Bengio, Y
err分享
err收藏
Reinforcement learning to adjust parametrized motor primitives to new situations
err2012-04-05
err121
PREAI
errKober, Jens; Wilhelm, Andreas; Oztop, Erhan; Peters, Jan
err分享
err收藏
Light Modulation of Enzyme Activity in Chloroplasts
err1976-02-01
err0
errOAAI
errLouise E. Anderson; Mordhay Avron
err分享
err收藏
Kernel methods in machine learning
err2008-06-01
err1.6K
errOAAI
errHofmann, Thomas; Schoelkopf, Bernhard; Smola, Alexander J.
err分享
err收藏
学者 查看更多内容