arrow
返回

Contrastive Learning Methods for Deep Reinforcement Learning

delete2023-01-01
delete3
delete
OA
AI
D
Di Wang *
M
Mengqi Hu
DOI:10.1109/ACCESS.2023.3312383delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Deep reinforcement learning (DRL) has shown promising performance in various application areas (e.g., games and autonomous vehicles). Experience replay buffer strategy and parallel learning strategy are widely used to boost the performances of offline and online deep reinforcement learning algorithms. However, state-action distribution shifts lead to bootstrap errors. Experience replay buffer learns policies with elder experience trajectories, limiting its application to off-policy algorithms. Balancing the new and the old experience is challenging. Parallel learning strategies can train policies with online experiences. However, parallel environmental instances organize the agent pool inefficiently with higher simulation or physical costs. To overcome these shortcomings, we develop four lightweight and effective DRL algorithms, instance-actor, parallel-actor, instance-critic, and parallel-critic methods, to contrast different-age trajectory experiences. We train the contrast DRL according to the received rewards and proposed contrast loss, which is calculated by designed positive/negative keys. Our benchmark experiments using PyBullet robotics environments show that our proposed algorithm matches or is better than the state-of-the-art DRL algorithms.
Keyword:
Contrastive learning
deep reinforcement learning
different-age experience
experience replay buffer
parallel learning

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

University of Illinois System 封面图
University of Illinois System
学者数:
6.8W
论文数: 6.2W
被引数: 644
引用论文

引用论文

Enhanced Off-Policy Reinforcement Learning With Focused Experience Replay
err2021-01-01
err7
errOAAI
errKong, Seung-Hyun; Nahrendra, I. Made Aswin; Paek, Dong-Hee
err分享
err收藏
Socioeconomic status affects the Oxford knee score and Short-Form 12 score following total knee replacement
err2013-01-01
err0
PREAI
errN. D. Clement; P. J. Jenkins; Y. X. Nie; J. T. Patton; S. J. Breusch; C. R. Howie; L. C. Biant
err分享
err收藏
Influence of SSTT, Ageing Regime and Stretching on IGC, Complex of Properties and Precipitation Behavior of 6013 Alloy
err2000-05-09
err0
PREAI
errV.G. Davydov; V.S. Siniavski; L.B. Ber; K.H. Rendigs; Gerhard Tempus; V.D. Valkov; V.D. Kalinin; Ye.V. Titkova; O.G. Ukolova; Ye.A. Lukina; Ye.I. Shvechkov; Ye.Ya. Kaputkin
err分享
err收藏
err分享
err收藏
Single-Image Reflection Removal Using Deep Learning: A Systematic Review
err2022-01-01
err13
errOAAI
errAmanlou, Ali; Suratgar, Amir Abolfazl; Tavoosi, Jafar; Mohammadzadeh, Ardashir; Mosavi, Amir
err分享
err收藏
Foundations and Modeling of Dynamic Networks Using Dynamic Graph Neural Networks: A Survey
err2021-01-01
err163
errOAAI
errSkarding, Joakim; Gabrys, Bogdan; Musial, Katarzyna
err分享
err收藏
A Robust Approach for Continuous Interactive Actor-Critic Algorithms
err2021-01-01
err14
errOAAI
errMillan-Arias, Cristian C.; Fernandes, Bruno J. T.; Cruz, Francisco; Dazeley, Richard; Fernandes, Sergio
err分享
err收藏
学者 查看更多内容