返回
Policy gradient in Lipschitz Markov Decision Processes
DOI:10.1007/s10994-015-5484-1.png)
摘要
En 中文
This paper is about the exploitation of Lipschitz continuity properties for Markov Decision Processes to safely speed up policy-gradient algorithms. Starting from assumptions about the Lipschitz continuity of the state-transition model, the reward function, and the policies considered in the learning process, we show that both the expected return of a policy and its gradient are Lipschitz continuous w.r.t. policy parameters. By leveraging such properties, we define policy-parameter updates that guarantee a performance improvement at each iteration. The proposed methods are empirically evaluated and compared to other related approaches using different configurations of three popular control scenarios: the linear quadratic regulator, the mass-spring-damper system and the ship-steering control.
Keyword:
Reinforcement learning
Markov Decision Process
Lipschitz continuity
Policy gradient algorithm
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
2.9
论文数:
2.7K
被引数:
3.4W
机构
引用论文
Calcium-Binding Capacity of Centrin2 Is Required for Linear POC5 Assembly but Not for Nucleotide Excision Repair
PLoS ONE
IF0

