Return
Stability-constrained policy optimization under unknown rewards
DOI:10.1016/j.ifacsc.2026.100366.png)
Abstract
En 中文
A major challenge in reinforcement learning (RL) is guaranteeing an agent's closed-loop stability under unknown, possibly sparse, reward functions. While model-free RL is flexible to a variety of systems and rewards, model-based control strategies such as optimization-based control naturally accommodate prior system models to provide guarantees on safety and stability. However, these models may not be representative of the true global performance objective, resulting in suboptimal policies. In this paper, we present a policy search RL approach that decouples the stability requirement from the global performance objective. The key idea is to use an optimization-based policy structure as an effective stabilizing parameterization with which the agent can learn to maximize an unknown reward in a model-free fashion. Specifically, the agent employs a predictive control architecture and implicitly learns a stabilizing terminal cost, which is constructed through fixed-point iterations of the discrete algebraic Riccati equation. By implicitly differentiating this fixed-point, derivatives of the stability condition inform policy gradients. The proposed approach is shown to design highperformance, stabilizing policies for various sparse, global performance objectives. Furthermore, the proposed approach can account for uncertainty in the dynamics using the stochastic discrete algebraic Riccati equation to promote robust stability. This work demonstrates a principled policy search RL approach, integrating prior models and system observations in an agent's design, towards safe and reliable decision-making under uncertainty. (c) 2026 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
Keywords:
Safe reinforcement learning
Stability-constrained policy search
Implicit differentiation
Predictive control
Journal
I
IF:
1.8
Papers:
80
Citations:
317

