arrow
返回

Deep reinforcement learning using least-squares truncated temporal-difference

delete2023-03-16
delete2
delete
OA
AI
J
Junkai Ren
Y
Yixing Lan
徐鑫 封面图
徐鑫 (Xin Xu)
Y
Yichuan Zhang
Q
Qiang Fang *
Y
Yujun Zeng
DOI:10.1049/cit2.12202delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Policy evaluation (PE) is a critical sub-problem in reinforcement learning, which estimates the value function for a given policy and can be used for policy improvement. However, there still exist some limitations in current PE methods, such as low sample efficiency and local convergence, especially on complex tasks. In this study, a novel PE algorithm called Least-Squares Truncated Temporal-Difference learning ((LSTD)-D-2) is proposed. In (LSTD)-D-2, an adaptive truncation mechanism is designed, which effectively takes advantage of the fast convergence property of Least-Squares Temporal Difference learning and the asymptotic convergence property of Temporal Difference learning (TD). Then, two feature pre-training methods are utilised to improve the approximation ability of (LSTD)-D-2. Furthermore, an Actor-Critic algorithm based on (LSTD)-D-2 and pre-trained feature representations (ACLPF) is proposed, where (LSTD)-D-2 is integrated into the critic network to improve learning-prediction efficiency. Comprehensive simulation studies were conducted on four robotic tasks, and the corresponding results illustrate the effectiveness of (LSTD)-D-2. The proposed ACLPF algorithm outperformed DQN, ACER and PPO in terms of sample efficiency and stability, which demonstrated that (LSTD)-D-2 can be applied to online learning control problems by incorporating it into the actor-critic architecture.
Keyword:
Deep reinforcement learning
policy evaluation
temporal difference
value function approximation
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

CAAI Transactions on Intelligence Technology 封面图
CAAI Transactions on Intelligence Technology
IF:
7.3
论文数:
669
被引数:
2.4K

机构

N
national university of defense technology - china
学者数:
1.8W
论文数: 1.4W
被引数: 9
引用论文

引用论文

Opening Platforms: How, When and Why?
err2008-01-01
err0
PREAI
errThomas R. Eisenmann; Geoffrey Parker; Marshall W. Van Alstyne
err分享
err收藏
err分享
err收藏
Evidence for a Ubiquitous Seismic Discontinuity at the Base of the Mantle
err1999-11-12
err0
PREAI
errIgor Sidorin; Michael Gurnis; Don V. Helmberger
err分享
err收藏
Social Science Concepts
err
IF0
err2019-10-17
err0
PREAI
errGary Goertz
err分享
err收藏
Predictors of COVID-19 Vaccine Uptake in Healthcare Workers
err2021-12-16
err0
errOAAI
errPetros Galanis; Ioannis Moisoglou; Irene Vraka; Olga Siskou; Olympia Konstantakopoulou; Aglaia Katsiroumpa; Daphne Kaitelidou
err分享
err收藏
学者 查看更多内容