arrow
Return

Asynchronous Parallel Policy Gradient Methods for the Linear Quadratic Regulator

delete2025-07-01
delete0
PRE
AI
F
Feiran Zhao
X
Xingyu Sha
游科友 (Keyou You)
DOI:10.1109/TAC.2025.3543128delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Learning policies in an asynchronous parallel way is essential to numerous successes of reinforcement learning for solving complex problems. However, their convergence has not been rigorously evaluated. To improve the theoretical understanding, we adopt the asynchronous parallel zero-order policy gradient (AZOPG) method to solve the continuous-time linear quadratic regulation problem. Specifically, multiple workers independently perform system rollouts to estimate zero-order policy gradients (PGs), which are then aggregated in a central node for policy updates. Moreover, each worker is allowed to interact with the central node <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">asynchronously</i>, leading to delayed PG estimates. By quantifying the convergence rate of AZOPG, we show a linear speedup property both in theory and simulation, which reveals the advantages of using asynchronous parallel workers in learning policies.
Keywords:
Asynchronous parallel methods
linear quadratic regulator (LQR)
linear system
policy gradient (PG)

Journal

IEEE Transactions on Automatic Control cover
IEEE Transactions on Automatic Control
IF:
7
Papers:
1.3W
Citations:
6.7W

Organization

T
tsinghua university
Scholars:
11.8W
Papers: 10.0W
Citations: 137