返回
Q-Learning for Continuous-Time Linear Systems: A Data-Driven Implementation of the Kleinman Algorithm
DOI:10.1109/TSMC.2022.3145693.png)
摘要
En 中文
A data-driven strategy to estimate the optimal feedback and the value function in an infinite-horizon, continuous-time, linear-quadratic optimal control problem for an unknown system is proposed. The method permits the construction of the optimal policy without any knowledge of the model, without requiring that the time derivatives of the state are available for the design, and without even assuming that an initial stabilizing feedback policy is available. Two alternative architectures are discussed: the first scheme revolves around the periodic computation of some matrix inversions involving the Q-function, whereas the second approach relies on a purely continuous-time implementation of some dynamic systems whose trajectories are uniformly attracted by the solutions to the above algebraic equations. Interestingly, the proposed strategy essentially constitutes a (direct) data-driven implementation of the celebrated Kleinman algorithm, hence subsuming the particularly appealing features of the latter, such as quadratic monotone convergence to the optimal solution. The theory is then validated by the means of practically motivated applications.
Keyword:
Optimal control
Trajectory
Costs
Convergence
Symmetric matrices
Riccati equations
Q-learning
Linear systems
optimal control
reinforcement learning
uncertain
unknown systems
期刊
IF:
10.5
论文数:
1.1W
被引数:
5.0W
机构
引用论文
Value iteration and adaptive dynamic programming for data-driven adaptive optimal control design数据驱动的自适应最优控制设计的值迭代和自适应动态规划
AUTOMATICA
IF5.9
Data-Based Adaptive Critic Designs for Nonlinear Robust Optimal Control With Uncertain Dynamics基于数据的具有不确定动态的非线性鲁棒最优控制的自适应评论家设计
Adaptive optimal control for continuous-time linear systems based on policy iteration基于策略迭代的连续时间线性系统自适应最优控制
AUTOMATICA
IF5.9

