Return
A basic formula for performance gradient estimation of semi-Markov decision processes
DOI:10.1016/j.ejor.2012.08.010.png)
Abstract
En 中文
This paper presents a basic formula for performance gradient estimation of semi-Markov decision processes (SMDPs) under average-reward criterion. This formula directly follows from a sensitivity equation in perturbation analysis. With this formula, we develop three sample-path-based gradient estimation algorithms by using a single sample path. These algorithms naturally extend many gradient estimation algorithms for discrete-time Markov systems to continuous time semi-Markov models. In particular, they require less storage than the algorithm in the literature. (C) 2012 Elsevier B.V. All rights reserved.
Keywords:
Markov processes
Semi-Markov decision processes
Sample-path-based gradient estimation
Perturbation analysis
Journal
IF:
6
Papers:
2.2W
Citations:
6.4W

