返回
Contrastive Learning Methods for Deep Reinforcement Learning
DOI:10.1109/ACCESS.2023.3312383.png)
摘要
En 中文
Deep reinforcement learning (DRL) has shown promising performance in various application areas (e.g., games and autonomous vehicles). Experience replay buffer strategy and parallel learning strategy are widely used to boost the performances of offline and online deep reinforcement learning algorithms. However, state-action distribution shifts lead to bootstrap errors. Experience replay buffer learns policies with elder experience trajectories, limiting its application to off-policy algorithms. Balancing the new and the old experience is challenging. Parallel learning strategies can train policies with online experiences. However, parallel environmental instances organize the agent pool inefficiently with higher simulation or physical costs. To overcome these shortcomings, we develop four lightweight and effective DRL algorithms, instance-actor, parallel-actor, instance-critic, and parallel-critic methods, to contrast different-age trajectory experiences. We train the contrast DRL according to the received rewards and proposed contrast loss, which is calculated by designed positive/negative keys. Our benchmark experiments using PyBullet robotics environments show that our proposed algorithm matches or is better than the state-of-the-art DRL algorithms.
Keyword:
Contrastive learning
deep reinforcement learning
different-age experience
experience replay buffer
parallel learning
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
A Survey of Domain-Specific Architectures for Reinforcement Learning用于强化学习的特定领域体系结构的调查
IEEE ACCESS
IF3.6
Sample Efficient Reinforcement Learning Method via High Efficient Episodic Memory基于高效情景记忆的样本高效强化学习方法
IEEE ACCESS
IF3.6
Foundations and Modeling of Dynamic Networks Using Dynamic Graph Neural Networks: A Survey
IEEE ACCESS
IF3.6

