arrow
Return

A basic formula for performance gradient estimation of semi-Markov decision processes

delete2013-01-01
delete12
PRE
AI
李彦杰 cover
李彦杰 (Yanjie Li) *
F
Fang Cao
DOI:10.1016/j.ejor.2012.08.010delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This paper presents a basic formula for performance gradient estimation of semi-Markov decision processes (SMDPs) under average-reward criterion. This formula directly follows from a sensitivity equation in perturbation analysis. With this formula, we develop three sample-path-based gradient estimation algorithms by using a single sample path. These algorithms naturally extend many gradient estimation algorithms for discrete-time Markov systems to continuous time semi-Markov models. In particular, they require less storage than the algorithm in the literature. (C) 2012 Elsevier B.V. All rights reserved.
Keywords:
Markov processes
Semi-Markov decision processes
Sample-path-based gradient estimation
Perturbation analysis

Journal

European Journal of Operational Research cover
European Journal of Operational Research
IF:
6
Papers:
2.2W
Citations:
6.4W

Organization

H
harbin institute of technology
Scholars:
8.0W
Papers: 6.6W
Citations: 66
B
Beijing Jiaotong University
Scholars:
2.2W
Papers: 1.7W
Citations: 1.2W