arrow
返回

Efficient Off-Policy Q-Learning for Data-Based Discrete-Time LQR Problems

delete2023-05-01
delete19
delete
OA
AI
V
Victor G. Lopez *
M
Mohammad Alsalti
M
Matthias A. Müller
DOI:10.1109/TAC.2023.3235967delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This article introduces and analyzes an improved Q-learning algorithm for discrete-time linear time-invariant systems. The proposed method does not require any knowledge of the system dynamics, and it enjoys significant efficiency advantages over other data-based optimal control methods in the literature. This algorithm can be fully executed offline, as it does not require to apply the current estimate of the optimal input to the system as in on-policy algorithms. It is shown that a PE input, defined from an easily tested matrix rank condition, guarantees the convergence of the algorithm. A data-based method is proposed to design the initial stabilizing feedback gain that the algorithm requires. Robustness of the algorithm in the presence of noisy measurements is analyzed. We compare the proposed algorithm in simulation to different direct and indirect data-based control design methods.
Keyword:
Q-learning
Heuristic algorithms
Data models
Convergence
Trajectory
Prediction algorithms
Linear systems
Data-based control
optimal control
reinforcement learning (RL)

期刊

IEEE Transactions on Automatic Control 封面图
IEEE Transactions on Automatic Control
IF:
7
论文数:
1.3W
被引数:
6.7W

机构

L
Leibniz University Hannover
学者数:
1.1W
论文数: 8.5K
被引数: 1.1W
引用论文

引用论文

The Steiner number of a graph
err2002-01-01
err0
errOAAI
errGary Chartrand; Ping Zhang
err分享
err收藏
340,000-Year Centennial-Scale Marine Record of Southern Hemisphere Climatic Oscillation
err2003-08-15
err0
PREAI
errKatharina Pahnke; Rainer Zahn; Henry Elderfield; Michael Schulz
err分享
err收藏
Cognitive Strategies for Special Education
err
IF0
err2017-09-13
err0
PREAI
errAdrian F. Ashman; Robert N.F. Conway
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容