返回
Mixed experience sampling for off-policy reinforcement learning
DOI:10.1016/j.eswa.2024.124017.png)
摘要
En 中文
In deep reinforcement learning, experience replay is usually used to improve data efficiency and alleviate experience forgetting. However, online reinforcement learning is often influenced by the index of experience, which usually makes the phenomenon of unbalanced sampling. In addition, most experience replay methods ignore the differences among experiences, and cannot make full use of all experiences. Especially many nearpolicy experiences relatively relevant to the current policy are wasted, despite of the fact that they are beneficial for improving sample efficiency. This paper theoretically analyzes the influence of various factors on experience sampling, and then proposes a sampling method for experience replay based on frequency and similarity (FSER) to alleviate unbalanced sampling and increase the value of the sampled experiences. FSER prefers experiences that are rarely sampled or highly relevant to the current policy. FSER plays a critical role to balance the experience forgetting and wasting problems. Finally, FSER is combined with TD3 to achieve the state-of-the-art results in multiple tasks.
Keyword:
Reinforcement learning
Experience replay
Experience sampling
Off-policy learning
Exploitation
期刊
IF:
7.5
论文数:
2.9W
被引数:
10.2W
机构
引用论文
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
Arachidonic Acid Metabolism by Human Cardiovascular CYP2J2 Is Modulated by Doxorubicin
Biochemistry
IF0

