arrow
返回

Fast online Q(λ)

delete1998-01-01
delete50
delete
OA
AI
W
Wierling, M *
J
Jürgen Schmidhuber
DOI:10.1023/A:1007562800292delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Q(lambda)-learning uses TD(lambda)-methods to accelerate Q-learning. The update complexity of previous online Q(lambda) implementations based on lookup tables is bounded by the size of the state/action space. Our faster algorithm's update complexity is bounded by the number of actions. The method is based on the observation that Q-value updates may be postponed until they are needed.
Keyword:
reinforcement learning
Q-learning
TD(lambda)
online Q(lambda)
lazy learning
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Machine Learning 封面图
Machine Learning
IF:
2.9
论文数:
2.7K
被引数:
3.4W

机构

暂无机构信息
引用论文

引用论文

暂无论文信息