Return
Policy-Adjustable Q-Learning for Data-Driven Nonlinear Optimal Tracking Control
DOI:10.1109/tnnls.2026.3672136.png)
Abstract
En 中文
This article investigates a novel policy-adjustable Q-learning (PA-QL) algorithm aimed at addressing the optimal tracking control (OTC) problem for nonlinear discrete-time (DT) systems with enhanced adaptability and flexibility. A novel iteration scheme is developed that integrates the control weights into the augmented neural network (NN) input, thereby reformulating the learning process to explicitly characterize the optimal policy as a function of the adjustable weights. Consequently, the learned control policy is not constrained by predetermined weights, allowing for dynamic adjustment after offline training is completed. Moreover, such adjustments can be performed online seamlessly, offering substantially greater flexibility in adapting to changes in operating conditions or control objectives. Finally, the effectiveness of the proposed algorithm is established through rigorous theoretical analysis and further validated by simulation studies.
Keywords:
Artificial neural networks
Q-learning
Optimal control
Training
Cost function
Vectors
Approximation algorithms
Switches
Data models
Computational modeling
Adaptive dynamic programming (ADP)
neural network (NN)
optimal tracking control (OTC)
policy adjustable
Journal
IF:
8.9
Papers:
7.6K
Citations:
7.2W
Organization
Cited Papers
No cited papers available

