Return
Non-stationary value iteration for adaptive average control of piecewise deterministic Markov processes
O
F
A
DOI:10.1016/j.nahs.2025.101622.png)
Abstract
En 中文
The main goal of this paper is to present a non-stationary value iteration scheme for the adaptive average control of Piecewise Deterministic Markov Processes (PDMPs), introduced by M.H.A. Davis in Davis (1984, 1993) as a family of continuous-time Markov processes punctuated by random jumps and with inter-jump movement driven by a deterministic flow. It is assumed in this paper that there are no boundary jumps. We study the adaptive average optimal control problem of PDMPs, considering that the jump intensity λ, the post-jump transition kernel Q, as well as the cost C depend on an unknown parameter β∗. For a sequence of strongly consistent estimators {βn∗} of β∗ (that is, βn∗ converge to β∗ almost surely) a non-stationary value iteration (depending on the current estimate βn∗) is shown to be optimal for the long-run average control problem. We assume a total variation norm condition on the parameters λ and Q of the process (which generalizes the minorization condition considered in Costa, Dufour and Genadot (2024), resulting in a span-contraction operator. The paper concludes with a numerical example.
Keywords:
PDMPs
adaptive average control
value iteration
non-stationary scheme
parameter estimation
Journal
N
IF:
4.1
Papers:
1.4K
Citations:
3.1K
