arrow
返回

Model-free LQR design by Q-function learning

delete2022-03-01
delete13
PRE
AI
M
Milad Farjadnasab
M
Maryam Babazadeh *
DOI:10.1016/j.automatica.2021.110060delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Reinforcement learning methods such as Q-learning have shown promising results in the model-free design of linear quadratic regulator (LQR) controllers for linear time-invariant (LTI) systems. However, challenges such as sample-efficiency, sensitivity to hyper-parameters, and compatibility with classical control paradigms limit the integration of such algorithms in critical control applications. This paper aims to take some steps towards bridging the well-known classical control requirements and learning algorithms by using optimization frameworks and properties of conic constraints. Accordingly, a new off-policy model-free approach is proposed for learning the Q-function and designing the discrete-time LQR controller. The design procedure is based on non-iterative semi-definite programs (SDP) with linear matrix inequality (LMI) constraints. It is sample-efficient, inherently robust to model uncertainties, and does not require an initial stabilizing controller. The proposed model-free approach is extended to distributed control of interconnected systems, as well. The performance of the presented design is evaluated on several stable and unstable synthetic systems. The data-driven control scheme is also implemented on the IEEE 39-bus New England power grid. The results confirm optimality, sample-efficiency, and satisfactory performance of the proposed approach in centralized and distributed design. (C) 2021 Elsevier Ltd. All rights reserved.
Keyword:
Convex optimization
Semi-definite programming (SDP)
Linear quadratic regulation (LQR)
Q-learning
Distributed control

期刊

Automatica 封面图
Automatica
IF:
5.9
论文数:
1.2W
被引数:
5.2W

机构

S
Sharif University of Technology
学者数:
1.1W
论文数: 1.1W
被引数: 9.5K
引用论文

引用论文

err分享
err收藏
err分享
err收藏
Distributed Adaptive Control of Synchronization in Complex Networks
err2012-08-01
err344
PREAI
errYu, Wenwu; DeLellis, Pietro; Chen, Guanrong; di Bernardo, Mario; Kurths, Juergen
err分享
err收藏
学者 查看更多内容