arrow
返回

Policy gradient in Lipschitz Markov Decision Processes

delete2015-03-03
delete59
delete
OA
AI
M
Matteo Pirotta *
M
Marcello Restelli
L
Luca Bascetta
DOI:10.1007/s10994-015-5484-1delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
This paper is about the exploitation of Lipschitz continuity properties for Markov Decision Processes to safely speed up policy-gradient algorithms. Starting from assumptions about the Lipschitz continuity of the state-transition model, the reward function, and the policies considered in the learning process, we show that both the expected return of a policy and its gradient are Lipschitz continuous w.r.t. policy parameters. By leveraging such properties, we define policy-parameter updates that guarantee a performance improvement at each iteration. The proposed methods are empirically evaluated and compared to other related approaches using different configurations of three popular control scenarios: the linear quadratic regulator, the mass-spring-damper system and the ship-steering control.
Keyword:
Reinforcement learning
Markov Decision Process
Lipschitz continuity
Policy gradient algorithm
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Machine Learning 封面图
Machine Learning
IF:
2.9
论文数:
2.7K
被引数:
3.4W

机构

P
Polytechnic University of Milan
学者数:
2.0W
论文数: 1.8W
被引数: 24
引用论文

引用论文

Natural Actor-Critic
err2008-03-01
err643
PREAI
errPeters, Jan; Schaal, Stefan
err分享
err收藏
Calcium-Binding Capacity of Centrin2 Is Required for Linear POC5 Assembly but Not for Nucleotide Excision Repair
err2013-07-02
err0
errOAAI
errTiago J. Dantas; Owen M. Daly; Pauline C. Conroy; Martin Tomas; Yifan Wang; Pierce Lalor; Peter Dockery; Elisa Ferrando-May; Ciaran G. Morrison
err分享
err收藏
NeuroGrid: recording action potentials from the surface of the brainNeuroGrid: 记录来自大脑表面的动作电位
err2014-12-22
err0
errOAAI
errDion Khodagholy; Jennifer N Gelinas; Thomas Thesen; Werner Doyle; Orrin Devinsky; George G Malliaras; György Buzsáki
err分享
err收藏
err分享
err收藏
Learning model-free robot control by a Monte Carlo EM algorithm
err2009-08-11
err44
errOAAI
errVlassis, Nikos; Toussaint, Marc; Kontes, Georgios; Piperidis, Savas
err分享
err收藏
Theoretical Study of Magnesium Fluoride in Aqueous Solution
err2011-08-17
err0
PREAI
errNaoto Shibata; Hirofumi Sato; Shigeyoshi Sakaki; Yuji Sugita
err分享
err收藏
A longitudinal study of long-term quality of life after ileal pouch-anal anastomosis
err2003-04-01
err0
PREAI
errRobert M Weinryb; Lars Liljeqvist; Bertil Poppen; J.Petter Gustavsson
err分享
err收藏
学者 查看更多内容