arrow
Return

Q-Learning for Continuous-Time Linear Systems: A Data-Driven Implementation of the Kleinman Algorithm

delete2022-10-01
delete11
PRE
AI
C
Corrado Possieri *
M
Mario Sassano
DOI:10.1109/TSMC.2022.3145693delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
A data-driven strategy to estimate the optimal feedback and the value function in an infinite-horizon, continuous-time, linear-quadratic optimal control problem for an unknown system is proposed. The method permits the construction of the optimal policy without any knowledge of the model, without requiring that the time derivatives of the state are available for the design, and without even assuming that an initial stabilizing feedback policy is available. Two alternative architectures are discussed: the first scheme revolves around the periodic computation of some matrix inversions involving the Q-function, whereas the second approach relies on a purely continuous-time implementation of some dynamic systems whose trajectories are uniformly attracted by the solutions to the above algebraic equations. Interestingly, the proposed strategy essentially constitutes a (direct) data-driven implementation of the celebrated Kleinman algorithm, hence subsuming the particularly appealing features of the latter, such as quadratic monotone convergence to the optimal solution. The theory is then validated by the means of practically motivated applications.
Keywords:
Optimal control
Trajectory
Costs
Convergence
Symmetric matrices
Riccati equations
Q-learning
Linear systems
optimal control
reinforcement learning
uncertain
unknown systems

Journal

IEEE Transactions on Cybernetics cover
IEEE Transactions on Cybernetics
IF:
10.5
Papers:
1.1W
Citations:
5.0W

Organization

C
consiglio nazionale delle ricerche (cnr)
Scholars:
6.2W
Papers: 5.7W
Citations: 48
U
University of Rome Tor Vergata
Scholars:
2.5W
Papers: 1.8W
Citations: 2.0W
Cited Papers

Cited Papers

Ion-specificity and surface water dynamics in protein solutions
err2018-01-01
err0
errOAAI
errTadeja Janc; Miha Lukšič; Vojko Vlachy; Baptiste Rigaud; Anne-Laure Rollet; Jean-Pierre Korb; Guillaume Mériguet; Natalie Malikova
errShare
errSave
errShare
errSave
A tutorial review of economic model predictive control methods
err2014-08-01
err587
PREAI
errEllis, Matthew; Durand, Helen; Christofides, Panagiotis D.
errShare
errSave
Adaptive optimal control for continuous-time linear systems based on policy iteration
err2009-02-01
err671
PREAI
errVrabie, D.; Pastravanu, O.; Abu-Khalaf, M.; Lewis, F. L.
errShare
errSave
CaCu3Ti4O12 single crystals: insights on growth and nanoscopic investigation
err2011-01-01
err0
errOAAI
errPatrick Fiorenza; Vito Raineri; Stefan G. Ebbinghaus; Raffaella Lo Nigro
errShare
errSave
Periprosthetic Infections
err1987-07-01
err0
PREAI
errDrogo K. Montague
errShare
errSave
researcher View more