返回
Importance sampling policy gradient algorithms in reproducing kernel Hilbert space
DOI:10.1007/s10462-017-9579-x.png)
摘要
En 中文
Modeling policies in reproducing kernel Hilbert space (RKHS) offers a very flexible and powerful new family of policy gradient algorithms called RKHS policy gradient algorithms. They are designed to optimize over a space of very high or infinite dimensional policies. As a matter of fact, they are known to suffer from a large variance problem. This critical issue comes from the fact that updating the current policy is based on a functional gradient that does not exploit all old episodes sampled by previous policies. In this paper, we introduce a generalized RKHS policy gradient algorithm that integrates the following important ideas: (i) policy modeling in RKHS; (ii) normalized importance sampling, which helps reduce the estimation variance by reusing previously sampled episodes in a principled way; and (iii) regularization terms, which avoid updating the policy too over-fit to sampled data. In the experiment section, we provide an analysis of the proposed algorithms through bench-marking domains. The experiment results show that the proposed algorithm can still enjoy a powerful policy modeling in RKHS and achieve more data-efficiency.
Keyword:
Reproducing kernel Hilbert space
Policy search
Reinforcement learning
Importance sampling
Policy gradient
Non-parametric
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
13.9
论文数:
6.1K
被引数:
1.9W
机构
引用论文
Real-time reinforcement learning by sequential Actor-Critics and experience replay
NEURAL NETWORKS
IF6.3
Reinforcement learning to adjust parametrized motor primitives to new situations
AUTONOMOUS ROBOTS
IF4.3

