arrow
Return

Constrained deterministic policy gradient adaptive dynamic programming using projected gradient update

delete2025-10-25
delete0
delete
OA
AI
T
Timotei Lala *
DOI:10.1080/21642583.2025.2546823delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Learning controllers for nonlinear complex systems using data-driven model-free adaptive dynamic programming (ADP) algorithms is a challenging problem. Traditional ADP methods often suffer from non-smooth convergence or divergence during learning, potentially leading to unstable control. In this paper, we propose a novel model-free constrained deterministic policy gradient adaptive dynamic programming (CDPGADP) for discrete-time general nonlinear systems. Our method constrains the difference between consecutive ADP iteration controllers using the method of projected gradient descent. Neural networks (NNs) are used to approximate both the state-action function and the controller. The constraint operates directly in the controller command space, effectively limiting the controller output during successive learning iterations. This algorithm is validated on a wind turbine pitch control task, learning to maintain a constant angular velocity at a prescribed nominal value under varying wind disturbances.
Keywords:
ADP
constrained ADP
deterministic policy gradient
projected gradient descent
Q-learning
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Systems Science and Control Engineering cover
Systems Science and Control Engineering
IF:
4.4
Papers:
486
Citations:
2.1K

Organization

P
Politehnica University of Timisoara
Scholars:
92
Papers: 52
Citations: 0