arrow
Return

Technical update: Least-squares temporal difference learning

delete2002-01-01
delete209
delete
OA
AI
B
Boyan, JA *
DOI:10.1023/A:1017936530646delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
TD(lambda) is a popular family of algorithms for approximate policy evaluation in large MDPs. TD(lambda) works by incrementally updating the value function after each observed transition. It has two major drawbacks: it may make inefficient use of data, and it requires the user to manually tune a stepsize schedule for good performance. For the case of linear value function approximations and lambda = 0, the Least-Squares TD (LSTD) algorithm of Bradtke and Barto (1996, Machine learning, 22:1-3, 33-57) eliminates all stepsize parameters and improves data efficiency. This paper updates Bradtke and Barto's work in three significant ways. First, it presents a simpler derivation of the LSTD algorithm. Second, it generalizes from lambda = 0 to arbitrary values of lambda; at the extreme of lambda = 1, the resulting new algorithm is shown to be a practical, incremental formulation of supervised linear regression. Third, it presents a novel and intuitive interpretation of LSTD as a model-based reinforcement learning technique.
Keywords:
reinforcement learning
temporal difference learning
value function approximation
linear least-squares methods
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Machine Learning cover
Machine Learning
IF:
2.9
Papers:
2.7K
Citations:
3.4W

Organization

No organization information available
Cited Papers

Cited Papers

No cited papers available