返回
Efficient Off-Policy Q-Learning for Data-Based Discrete-Time LQR Problems
DOI:10.1109/TAC.2023.3235967.png)
摘要
En 中文
This article introduces and analyzes an improved Q-learning algorithm for discrete-time linear time-invariant systems. The proposed method does not require any knowledge of the system dynamics, and it enjoys significant efficiency advantages over other data-based optimal control methods in the literature. This algorithm can be fully executed offline, as it does not require to apply the current estimate of the optimal input to the system as in on-policy algorithms. It is shown that a PE input, defined from an easily tested matrix rank condition, guarantees the convergence of the algorithm. A data-based method is proposed to design the initial stabilizing feedback gain that the algorithm requires. Robustness of the algorithm in the presence of noisy measurements is analyzed. We compare the proposed algorithm in simulation to different direct and indirect data-based control design methods.
Keyword:
Q-learning
Heuristic algorithms
Data models
Convergence
Trajectory
Prediction algorithms
Linear systems
Data-based control
optimal control
reinforcement learning (RL)
期刊
IF:
7
论文数:
1.3W
被引数:
6.7W
机构
引用论文
Data-Driven Finite-Horizon Approximate Optimal Control for Discrete-Time Nonlinear Systems Using Iterative HDP Approach使用迭代HDP方法对离散时间非线性系统进行数据驱动的有限水平近似最优控制
H∞ Tracking Control of Completely Unknown Continuous-Time Systems via Off-Policy Reinforcement Learning基于非策略强化学习的完全未知连续时间系统的h ∞ 跟踪控制
Adaptive Optimal Control of Unknown Constrained-Input Systems Using Policy Iteration and Neural Networks基于策略迭代和神经网络的未知约束输入系统的自适应最优控制

