arrow
返回

Residual Sarsa algorithm with function approximation

delete2017-11-10
delete1
PRE
AI
Q
Qiming Fu
H
Hu Wen
Q
Quan Liu
罗恒 (Heng Luo)
L
Lingyao Hu
J
Jianping Chen *
DOI:10.1007/s10586-017-1303-8delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In this work, we proposed an efficient algorithm named the residual Sarsa algorithm with function approximation (FARS) to improve the performance of the traditional Sarsa algorithm, and we use the gradient-descent method to update the function parameter vector. In the learning process, the Bellman residual method is adopted to guarantee the convergence of the algorithm, and a new rule for updating vectors of action-value functions is adopted to solve unstable and slow convergence problems. To accelerate the convergence rate of the algorithm, we introduce a new factor, named the forgotten factor, which can help improve the robustness of the algorithm's performance. Based on two classical reinforcement learning benchmark problems, the experimental results show that the FARS algorithm has better performance than other related reinforcement learning algorithms.
Keyword:
Reinforcement learning
Sarsa algorithm
Function approximation
Gradient descent
Bellman residual
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

C
Cluster Computing-The Journal of Networks Software Tools and Applications
IF:
4.1
论文数:
5.0K
被引数:
7.5K

机构

S
suzhou university of science & technology
学者数:
5.0K
论文数: 4.8K
被引数: 4
S
soochow university - china
学者数:
5.2W
论文数: 3.6W
被引数: 82
引用论文

引用论文

Evidence for a Ubiquitous Seismic Discontinuity at the Base of the Mantle
err1999-11-12
err0
PREAI
errIgor Sidorin; Michael Gurnis; Don V. Helmberger
err分享
err收藏
Reinforcement Learning for Energy Harvesting Point-to-Point Communications
err2016-05-01
err43
PREAI
errOrtiz, Andrea; Al-Shatri, Hussein; Li, Xiang; Weber, Tobias; Klein, Anja
err分享
err收藏
Case Study: 3D Modelling and Printing of a Plastic Respirator in Laboratory Conditions
err2021-12-23
err0
errOAAI
errMiriam Pekarcikova; Peter Trebuna; Marek Kliment; Stefan Kral
err分享
err收藏
Convergence results for single-step on-policy reinforcement-learning algorithms
err2000-01-01
err464
errOAAI
errSingh, S; Jaakkola, T; Littman, ML; Szepesvári, C
err分享
err收藏
学者 查看更多内容