arrow
返回

Sampling diversity driven exploration with state difference guidance

delete2022-10-01
delete3
PRE
AI
S
Shuai D. Han
S
Shuai Lü *
M
Meng Kang
J
Junwei Zhang
DOI:10.1016/j.eswa.2022.117418delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Exploration is one of the key issues of deep reinforcement learning, especially in the environments with sparse or deceptive rewards. Exploration based on intrinsic rewards can handle these environments. However, these methods cannot take both global interaction dynamics and local environment changes into account simultaneously. In this paper, we propose a novel intrinsic reward for off-policy learning, which not only encourages the agent to take actions not fully learned from a global perspective, but also instructs the agent to trigger remarkable changes in the environment from a local perspective. Meanwhile, we propose the doubleactors-double-critics framework to combine intrinsic rewards with extrinsic rewards to avoid the inappropriate combination of intrinsic and extrinsic rewards in previous methods. This framework can be applied to off policy learning algorithms based on the actor-critic method. We provide a comprehensive evaluation of our approach on the MuJoCo benchmark environments. The results demonstrate that our method can perform effective exploration in the environments with dense, deceptive and sparse rewards. Besides, we conduct sufficient ablation and quantitative analyses to intrinsic rewards. Furthermore, we also verify the superiority and rationality of our double-actors-double-critics framework through comparative experiments.
Keyword:
Reinforcement learning
Exploration
Intrinsic rewards
Off-policy
Actor-critic algorithm

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
2.9W
被引数:
10.2W

机构

J
Jilin University
学者数:
8.7W
论文数: 5.6W
被引数: 8.9K
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
Surface modified electrospun porous magnetic hollow fibers using secondary downstream collection solvent contouring
err2017-10-01
err0
errOAAI
errShuting Wu; Baolin Wang; Zeeshan Ahmad; Jie Huang; Ming-Wei Chang; Jing-Song Li
err分享
err收藏
Influencing factors on the accuracy of local geoid model
err2019-11-01
err0
errOAAI
errShazad Jamal Jalal; Tajul Ariffin Musa; Ami Hassan Md Din; Wan Anom Wan Aris; WenBin Shen; Muhammad Faiz Pa'suya
err分享
err收藏
Diversity-augmented intrinsic motivation for deep reinforcement learning
err2022-01-01
err13
PREAI
errDai, Tianhong; Du, Yali; Fang, Meng; Bharath, Anil Anthony
err分享
err收藏