Return
Constrained deterministic policy gradient adaptive dynamic programming using projected gradient update
DOI:10.1080/21642583.2025.2546823.png)
Abstract
En 中文
Learning controllers for nonlinear complex systems using data-driven model-free adaptive dynamic programming (ADP) algorithms is a challenging problem. Traditional ADP methods often suffer from non-smooth convergence or divergence during learning, potentially leading to unstable control. In this paper, we propose a novel model-free constrained deterministic policy gradient adaptive dynamic programming (CDPGADP) for discrete-time general nonlinear systems. Our method constrains the difference between consecutive ADP iteration controllers using the method of projected gradient descent. Neural networks (NNs) are used to approximate both the state-action function and the controller. The constraint operates directly in the controller command space, effectively limiting the controller output during successive learning iterations. This algorithm is validated on a wind turbine pitch control task, learning to maintain a constant angular velocity at a prescribed nominal value under varying wind disturbances.
Keywords:
ADP
constrained ADP
deterministic policy gradient
projected gradient descent
Q-learning
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
4.4
Papers:
486
Citations:
2.1K

