arrow
Return

An Actor-Critic Algorithm With Second-Order Actor and Critic

delete2017-06-01
delete9
delete
OA
AI
王静 cover
王静 (Jing Wang) *
I
Ioannis Ch. Paschalidis
DOI:10.1109/TAC.2016.2616384delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Actor-critic algorithms solve dynamic decision making problems by optimizing a performance metric of interest over a user-specified parametric class of policies. They employ a combination of an actor, making policy improvement steps, and a critic, computing policy improvement directions. Many existing algorithms use a steepest ascent method to improve the policy, which is known to suffer from slow convergence for ill-conditioned problems. In this paper, we first develop an estimate of the (Hessian) matrix containing the second derivatives of the performance metric with respect to policy parameters. Using this estimate, we introduce a new second-order policy improvement method and couple it with a critic using a second-order learning method. We establish almost sure convergence of the new method to a neighborhood of a policy parameter stationary point. We compare the new algorithm with some existing algorithms in two applications and demonstrate that it leads to significantly faster convergence.
Keywords:
Actor-critic algorithms
Markov decision processes
Newton's method
robotics
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Automatic Control cover
IEEE Transactions on Automatic Control
IF:
7
Papers:
1.3W
Citations:
6.7W

Organization

B
boston university
Scholars:
3.7W
Papers: 3.2W
Citations: 67