Return
(Approximate) iterated successive approximations algorithm for sequential decision processes
DOI:10.1007/s10479-012-1073-x.png)
Abstract
En 中文
The paper proves the convergence of (Approximate) Iterated Successive Approximations Algorithm for solving infinite-horizon sequential decision processes satisfying the monotone contraction assumption. At every stage of this algorithm, the value function at hand is used as a terminal reward to determine an (approximately) optimal policy for the one-period problem. This policy is then iterated for a (finite or infinite) number of times and the resulting return function is used as the starting value function for the next stage of the scheme. This method generalizes the standard successive approximations, policy iteration and Denardo's generalization of the latter.
Keywords:
Sequential decision processes
Markov decision chains
Successive approximations
Modified policy iteration
Journal
IF:
4.5
Papers:
8.0K
Citations:
2.1W
Organization
Cited Papers
A Novel Feature Identification Method of Pipeline In-Line Inspected Bending Strain Based on Optimized Deep Belief Network Model
Energies
IF0

