arrow
Return

Implicit incremental natural actor critic algorithm

delete2019-01-01
delete3
delete
OA
AI
R
Ryo Iwaki *
M
Minoru Asada
DOI:10.1016/j.neunet.2018.10.007delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Natural policy gradient (NPG) methods are promising approaches to finding locally optimal policy parameters. The NPG approach works well in optimizing complex policies with high-dimensional parameters, and the effectiveness of NPG methods has been demonstrated in many fields. However, the incremental estimation of the NPG is computationally unstable owing to its high sensitivity to the step-sizes values, especially to the one used to update the estimate of NPG. In this study, we propose a new incremental and stable algorithm for the NPG estimation. We call the proposed algorithm the implicit incremental natural actor critic (I2NAC), and it is based on the idea of the implicit update. The convergence analysis for I2NAC is provided. Theoretical analysis results indicate the stability of I2NAC and the instability of conventional incremental NPG methods. Numerical experiments were performed, and the results show that I2NAC is less sensitive to the values of the meta-parameters, including the step-size for the NPG update, compared to the existing incremental NPG method. (C) 2018 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY-NC-ND license.
Keywords:
Reinforcement learning
Natural policy gradient
Natural actor critic
Incremental learning
Implicit update
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Neural Networks cover
Neural Networks
IF:
6.3
Papers:
7.8K
Citations:
3.0W

Organization

O
osaka university
Scholars:
2.6W
Papers: 1.9W
Citations: 30