arrow
Return

Stochastic cubic-regularized policy gradient method

delete2022-11-01
delete3
PRE
AI
P
Pengfei Wang
H
Hongyu Wang
郑能干 (Nenggan Zheng) *
DOI:10.1016/j.knosys.2022.109687delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Policy-based reinforcement learning methods have achieved great achievements in real-world decisionmaking problems. However, the theoretical understanding of policy-based methods is still limited. Specifically, existing works mainly focus on first-order stationary point policies (FOSPs); in some very special reinforcement learning settings (e.g., tabular case and function approximation with restricted parametric policy classes) some works consider globally optimal policy. It is well-known that FOSPs could be undesirable local optima or saddle points, and obtaining a global optimum is generally NP-hard. In this paper, we propose a policy gradient method that provably converges to secondorder stationary point policies (SOSPs) for any differentiable policy classes. The proposed method is computationally efficient, and it judiciously uses cubic-regularized subroutines to escape saddle points while at the same time minimizing the Hessian-based computations. We prove that the method enjoys the sample complexity of O tilde (epsilon-3.5), which improves upon the current optimal complexity O tilde (epsilon-4.5). Finally, experimental results are provided to demonstrate the effectiveness of the method.(c) 2022 Elsevier B.V. All rights reserved.
Keywords:
Reinforcement learning
Policy gradient
Stochastic optimization
Non -convex optimization
Second-order stationary point

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

Z
zhejiang university
Scholars:
17.5W
Papers: 12.0W
Citations: 152