arrow
Return

Stability-constrained Markov Decision Processes using MPC

delete2022-09-01
delete5
delete
OA
AI
M
Mario Zanon *
S
Sébastien Gros
M
Michele Palladino
DOI:10.1016/j.automatica.2022.110399delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
In this paper, we consider solving discounted Markov Decision Processes (MDPs) under the constraint that the resulting policy is stabilizing. In practice MDPs are solved based on some form of policy approximation. We will leverage recent results proposing to use Model Predictive Control (MPC) as a structured approximator in the context of Reinforcement Learning, which makes it possible to introduce stability requirements directly inside the MPC-based policy. This will restrict the solution of the MDP to stabilizing policies by construction. Because the stability theory for MPC is most mature for the undiscounted MPC case, we will first show in this paper that stable discounted MDPs can be reformulated as undiscounted ones. This observation will entail that the undiscounted MPC-based policy with stability guarantees will produce the optimal policy for the discounted MDP if it is stable, and the best stabilizing policy otherwise. (C) 2022 Elsevier Ltd. All rights reserved.
Keywords:
Markov Decision Processes
Model Predictive Control
Stability
Safe reinforcement learning
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Automatica cover
Automatica
IF:
5.9
Papers:
1.2W
Citations:
5.2W

Organization

University of LAquila cover
University of LAquila
Scholars:
7.4K
Papers: 6.6K
Citations: 6.7K
I
IMT School for Advanced Studies Lucca
Scholars:
676
Papers: 701
Citations: 693