返回
Linear Quadratic Control Using Model-Free Reinforcement Learning
DOI:10.1109/TAC.2022.3145632.png)
摘要
En 中文
In this article, we consider linear quadratic (LQ) control problem with process and measurement noises. We analyze the LQ problem in terms of the average cost and the structure of the value function. We assume that the dynamics of the linear system is unknown and only noisy measurements of the state variable are available. Using noisy measurements of the state variable, we propose two model-free iterative algorithms to solve the LQ problem. The proposed algorithms are variants of policy iteration routine where the policy is greedy with respect to the average of all previous iterations. We rigorously analyze the properties of the proposed algorithms, including stability of the generated controllers and convergence. We analyze the effect of measurement noise on the performance of the proposed algorithms, the classical off-policy, and the classical Q-learning routines. We also investigate a model-building approach, inspired by adaptive control, where a model of the dynamical system is estimated and the optimal control problem is solved assuming that the estimated model is the true model. We use a benchmark to evaluate and compare our proposed algorithms with the classical off-policy, the classical Q-learning, and the policy gradient. We show that our model-building approach performs nearly identical to the analytical solution and our proposed policy iteration-based algorithms outperform the classical off-policy and the classical Q-learning algorithms on this benchmark but do not outperform the model-building approach.
Keyword:
Noise measurement
Costs
Dynamical systems
Adaptation models
Heuristic algorithms
Process control
Optimal control
Linear quadratic (LQ) control
reinforcement learning (RL)
期刊
IF:
7
论文数:
1.3W
被引数:
6.7W
机构
引用论文
Distribution of Substance P and Calcitonin Gene-related Peptide-like lmmunoreactive Nerve Fibers in the Rat Temporomandibular JointP物质和降钙素基因相关肽样免疫反应性神经纤维在大鼠颞下颌关节中的分布
Reinforcement learning for a class of continuous-time input constrained optimal control problems一类连续时间输入约束最优控制问题的强化学习
AUTOMATICA
IF5.9
Tumor acidosis enhances cytotoxic effects and autophagy inhibition by salinomycin on cancer cell lines and cancer stem cells
Oncotarget
IF0
Linear Quadratic Tracking Control of Partially-Unknown Continuous-Time Systems Using Reinforcement Learning基于强化学习的部分未知连续时间系统的线性二次跟踪控制

