返回
Value Iteration for Continuous-Time Linear Time-Invariant Systems
DOI:10.1109/TAC.2022.3169688.png)
摘要
En 中文
Two data-driven strategies for value iteration in linear quadratic optimal control problems over an infinite horizon are proposed. The two architectures share common features, since they both consist of a purely continuous-time control architecture and are based on the forward integration of the differential Riccati equation (DRE). They profoundly differ, instead, in the estimation mechanism of the vector field of the underlying DRE from collected data: The first relies on a characterization of properties of the advantage function associated to the problem, whereas the second is inspired by tools from adaptive control theory and ensures semi-global exponential convergence to the optimal solution. Advantages and drawbacks of the architectures are discussed, while the performance is validated via a benchmark numerical example.
Keyword:
Optimal control
Costs
Riccati equations
Reinforcement learning
Convergence
Adaptive control
Trajectory
linear systems
optimal control
reinforcement learning
期刊
IF:
7
论文数:
1.3W
被引数:
6.7W
机构
引用论文
Adaptive optimal control for continuous-time linear systems based on policy iteration基于策略迭代的连续时间线性系统自适应最优控制
AUTOMATICA
IF5.9
A process dissociation framework: Separating automatic from intentional uses of memory过程分离框架: 将自动与有意使用的内存分开
Integral Q-learning and explorized policy iteration for adaptive optimal control of continuous-time linear systems连续时间线性系统自适应最优控制的积分Q学习和探索性策略迭代
AUTOMATICA
IF5.9
没有更多内容

