arrow
返回

Recursive Least-Squares Temporal Difference With Gradient Correction

delete2021-08-01
delete2
PRE
AI
T
Tianheng Song
D
Dazi Li *
杨卫民 (Weimin Yang)
K
Kotaro Hirasawa
DOI:10.1109/TCYB.2019.2902342delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Since the late 1980s, temporal difference (TD) learning has dominated the research area of policy evaluation algorithms. However, the demand for the avoidance of TD defects, such as low data-efficiency and divergence in off-policy learning, has inspired the studies of a large number of novel TD-based approaches. Gradient-based and least-squares-based algorithms comprise the major part of these new approaches. This paper aims to combine advantages of these two categories to derive an efficient policy evaluation algorithm with O(n(2)) per-time-step runtime complexity. The least-squares-based framework is adopted, and the gradient correction is used to improve convergence performance. This paper begins with the revision of a previous O(n(3)) batch algorithm, least-squares TD with a gradient correction (LS-TDC) to regularize the parameter vector. Based on the recursive least-squares technique, an O(n(2)) counterpart of LS-TDC called RC is proposed. To increase data efficiency, we generalize RC with eligibility traces. An off-policy extension is also proposed based on importance sampling. In addition, the convergence analysis for RC as well as LS-TDC is given. The empirical results in both on-policy and off-policy benchmarks show that RC has a higher estimation accuracy than that of RLSTD and a significantly lower runtime complexity than that of LSTDC.
Keyword:
Policy evaluation
reinforcement learning (RL)
temporal differences (TDs)
value function approximation
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Cybernetics 封面图
IEEE Transactions on Cybernetics
IF:
10.5
论文数:
1.1W
被引数:
5.0W

机构

B
Beijing University of Chemical Technology
学者数:
3.1W
论文数: 2.2W
被引数: 4.5W
引用论文

引用论文

Viral pathogenesis of SARS-CoV-2 infection and male reproductive health
err2021-01-20
err0
errOAAI
errShubhadeep Roychoudhury; Anandan Das; Niraj Kumar Jha; Kavindra Kumar Kesari; Shatabhisha Roychoudhury; Saurabh Kumar Jha; Raghavender Kosgi; Arun Paul Choudhury; Norbert Lukac; Nithar Ranjan Madhu; Dhruv Kumar; Petr Slama
err分享
err收藏
Opening Platforms: How, When and Why?
err2008-01-01
err0
PREAI
errThomas R. Eisenmann; Geoffrey Parker; Marshall W. Van Alstyne
err分享
err收藏
err分享
err收藏
Evidence for a Ubiquitous Seismic Discontinuity at the Base of the Mantle
err1999-11-12
err0
PREAI
errIgor Sidorin; Michael Gurnis; Don V. Helmberger
err分享
err收藏
err分享
err收藏
Reinforcement Learning for Port-Hamiltonian Systems
err2015-05-01
err31
errOAAI
errSprangers, Olivier; Babuska, Robert; Nageshrao, Subramanya P.; Lopes, Gabriel A. D.
err分享
err收藏
学者 查看更多内容