arrow
Return

A switching control strategy for policy selection in stochastic Dynamic Programming problems☆

delete2025-01-01
delete0
PRE
AI
M
Massimo Tipaldi
R
Raffaele Iervolino *
P
Paolo Roberto Massenio
D
David Naso
DOI:10.1016/j.automatica.2024.111884delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This paper presents a switching control strategy as a criterion for policy selection in stochastic Dynamic Programming problems over an infinite time horizon. In particular, the Bellman operator, applied iteratively to solve such problems, is generalized to the case of stochastic policies, and formulated as a discrete-time switched affine system. Then, a Lyapunov-based policy selection strategy is designed to ensure the practical convergence of the resulting closed-loop system trajectories towards an appropriately chosen reference value function. This way, it is possible to verify how the chosen reference value function can be approached by using a stabilizing switching signal, the latter defined on a given finite set of stationary stochastic policies. Finally, the presented method is applied to the Value Iteration algorithm, and an illustrative example of a recycling robot is provided to demonstrate its effectiveness in terms of convergence performance. (c) 2024 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
Keywords:
Switched affine systems
Dynamic programming
Lyapunov function
Bellman operator
Stationary stochastic policies

Journal

Automatica cover
Automatica
IF:
5.9
Papers:
1.2W
Citations:
5.2W

Organization

P
Politecnico di Bari
Scholars:
4.0K
Papers: 4.0K
Citations: 6
U
University of Naples Federico II
Scholars:
4.7W
Papers: 3.6W
Citations: 51