返回
Model-free LQR design by Q-function learning
DOI:10.1016/j.automatica.2021.110060.png)
摘要
En 中文
Reinforcement learning methods such as Q-learning have shown promising results in the model-free design of linear quadratic regulator (LQR) controllers for linear time-invariant (LTI) systems. However, challenges such as sample-efficiency, sensitivity to hyper-parameters, and compatibility with classical control paradigms limit the integration of such algorithms in critical control applications. This paper aims to take some steps towards bridging the well-known classical control requirements and learning algorithms by using optimization frameworks and properties of conic constraints. Accordingly, a new off-policy model-free approach is proposed for learning the Q-function and designing the discrete-time LQR controller. The design procedure is based on non-iterative semi-definite programs (SDP) with linear matrix inequality (LMI) constraints. It is sample-efficient, inherently robust to model uncertainties, and does not require an initial stabilizing controller. The proposed model-free approach is extended to distributed control of interconnected systems, as well. The performance of the presented design is evaluated on several stable and unstable synthetic systems. The data-driven control scheme is also implemented on the IEEE 39-bus New England power grid. The results confirm optimality, sample-efficiency, and satisfactory performance of the proposed approach in centralized and distributed design. (C) 2021 Elsevier Ltd. All rights reserved.
Keyword:
Convex optimization
Semi-definite programming (SDP)
Linear quadratic regulation (LQR)
Q-learning
Distributed control
期刊
IF:
5.9
论文数:
1.2W
被引数:
5.2W
机构
引用论文
Linear matrix inequalities, Riccati equations, and indefinite stochastic linear quadratic controls线性矩阵不等式、Riccati方程和不定随机线性二次控制
Reinforcement learning for control: Performance, stability, and deep approximators用于控制的强化学习: 性能、稳定性和深度逼近

