arrow
返回

Correlation minimizing replay memory in temporal-difference reinforcement learning

delete2020-06-01
delete12
PRE
AI
M
Mirza Ramičić *
A
Andrea Bonarini
DOI:10.1016/j.neucom.2020.02.004delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Online reinforcement learning agents are now able to process an increasing amount of data which makes their approximation and compression into value functions a more demanding task. To improve approximation, thus the learning process itself, it has been proposed to select randomly a mini-batch of the past experiences that are stored in the replay memory buffer to be replayed at each learning step. In this work, we present an algorithm that classifies and samples the experiences into separate contextual memory buffers using an unsupervised learning technique. This allows each new experience to be associated to a mini-batch of the past experiences that are not from the same contextual buffer as the current one, thus further reducing the correlation between experiences. Experimental results show that the correlation minimizing sampling improves over Q-learning algorithms with uniform sampling, and that a significant improvement can be observed when coupled with the sampling methods that prioritize on the experience temporal difference error. (C) 2020 Elsevier B.V. All rights reserved.
Keyword:
Reinforcement learning
Temporal-difference learning
Replay memory
Artificial neural networks

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

P
Polytechnic University of Milan
学者数:
2.0W
论文数: 1.8W
被引数: 24
C
czech technical university prague
学者数:
6.6K
论文数: 5.3K
被引数: 3
引用论文

引用论文

err分享
err收藏
Multitask learning多任务学习
err1997-01-01
err4.9K
errOAAI
errCaruana, R
err分享
err收藏
err分享
err收藏