arrow
Return

Policy Iteration-Based Learning Design for Linear Continuous-Time Systems Under Initial Stabilizing OPFB Policy

delete2024-11-01
delete1
PRE
AI
C
Chengye Zhang
C
Ci Chen *
F
Frank L. Lewis
S
Shengli Xie
DOI:10.1109/TCYB.2024.3418190delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Policy iteration (PI), an iterative method in reinforcement learning, has the merit of interactions with a little-known environment to learn a decision law through policy evaluation and improvement. However, the existing PI-based results for output-feedback (OPFB) continuous-time systems relied heavily on an initial stabilizing full state-feedback (FSFB) policy. It thus raises the question of violating the OPFB principle. This article addresses such a question and establishes the PI under an initial stabilizing OPFB policy. We prove that an off-policy Bellman equation can transform any OPFB policy into an FSFB policy. Based on this transformation property, we revise the traditional PI by appending an additional iteration, which turns out to be efficient in approximating the optimal control under the initial OPFB policy. We show the effectiveness of the proposed learning methods through theoretical analysis and a case study.
Keywords:
Control design
Optimal control
Regulation
Standards
Mathematical models
Convergence
Trajectory
Initial policy
output feedback (OPFB)
policy iteration (PI)
reinforcement learning (RL)

Journal

IEEE Transactions on Cybernetics cover
IEEE Transactions on Cybernetics
IF:
10.5
Papers:
1.1W
Citations:
5.0W

Organization

G
guangdong university of technology
Scholars:
3.0W
Papers: 2.0W
Citations: 36