Return
Average cost temporal-difference learning
DOI:10.1016/S0005-1098(99)00099-0.png)
Abstract
En 中文
We propose a variant of temporal-difference learning that approximates average and differential costs of an irreducible aperiodic Markov chain. Approximations are comprised of linear combinations of fixed basis functions whose weights are incrementally updated during a single endless trajectory of the Markov chain. We present a proof of convergence (with probability 1) and a characterization of the limit of convergence. We also provide a bound on the resulting approximation error that exhibits an interesting dependence on the mixing time of the Markov chain. The results parallel previous work by the authors, involving approximations of discounted cost-to-go. (C) 1999 Elsevier Science Ltd. All rights reserved.
Keywords:
dynamic programming
learning
average cost
reinforcement learning
neuro-dynamic programming
approximation
temporal differences
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

