Return
Asynchronous Parallel Policy Gradient Methods for the Linear Quadratic Regulator
DOI:10.1109/TAC.2025.3543128.png)
Abstract
En 中文
Learning policies in an asynchronous parallel way is essential to numerous successes of reinforcement learning for solving complex problems. However, their convergence has not been rigorously evaluated. To improve the theoretical understanding, we adopt the asynchronous parallel zero-order policy gradient (AZOPG) method to solve the continuous-time linear quadratic regulation problem. Specifically, multiple workers independently perform system rollouts to estimate zero-order policy gradients (PGs), which are then aggregated in a central node for policy updates. Moreover, each worker is allowed to interact with the central node <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">asynchronously</i>, leading to delayed PG estimates. By quantifying the convergence rate of AZOPG, we show a linear speedup property both in theory and simulation, which reveals the advantages of using asynchronous parallel workers in learning policies.
Keywords:
Asynchronous parallel methods
linear quadratic regulator (LQR)
linear system
policy gradient (PG)
Journal
IF:
7
Papers:
1.3W
Citations:
6.7W

