arrow
Return

Compatible natural gradient policy search

delete2019-05-20
delete8
delete
OA
AI
J
Joni Pajarinen *
H
Hong Linh Thai
R
Riad Akrour
J
Jan Peters
G
Gerhard Neumann
DOI:10.1007/s10994-019-05807-0delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Trust-region methods have yielded state-of-the-art results in policy search. A common approach is to use KL-divergence to bound the region of trust resulting in a natural gradient policy update. We show that the natural gradient and trust region optimization are equivalent if we use the natural parameterization of a standard exponential policy distribution in combination with compatible value function approximation. Moreover, we show that standard natural gradient updates may reduce the entropy of the policy according to a wrong schedule leading to premature convergence. To control entropy reduction we introduce a new policy search method called compatible policy search (COPOS) which bounds entropy loss. The experimental results show that COPOS yields state-of-the-art results in challenging continuous control tasks and in discrete partially observable tasks.
Keywords:
Reinforcement learning
Policy search
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Machine Learning cover
Machine Learning
IF:
2.9
Papers:
2.6K
Citations:
3.4W

Organization

U
University of Lincoln
Scholars:
2.6K
Papers: 2.5K
Citations: 3.9K
T
Technical University of Darmstadt
Scholars:
1.3W
Papers: 10.0K
Citations: 1.2W