arrow
Return

Stochastic Integrated ActorCritic for Deep Reinforcement Learning

delete2024-05-01
delete4
PRE
AI
J
Jiaohao Zheng
M
Mehmet Necip Kurt
X
Xiaodong Wang *
DOI:10.1109/TNNLS.2022.3212273delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We propose a deep stochastic actor-critic algorithm with an integrated network architecture and fewer parameters. We address stabilization of the learning procedure via an adaptive objective to the critic's loss and a smaller learning rate for the shared parameters between the actor and the critic. Moreover, we propose a mixed on-off policy exploration strategy to speed up learning. Experiments illustrate that our algorithm reduces the sample complexity by 50%-93% compared with the state-of-the-art deep reinforcement learning (RL) algorithms twin delayed deep deterministic policy gradient (TD3), soft actor-critic (SAC), proximal policy optimization (PPO), advantage actor-critic (A2C), and interpolated policy gradient (IPG) over continuous control tasks LunarLander, BipedalWalker, BipedalWalkerHardCore, Ant, and Minitaur in the OpenAI Gym.
Keywords:
Training
Task analysis
Complexity theory
Linear programming
Network architecture
Decoding
Tensors
Actor-critic
adaptive objective
deep reinforcement learning (RL)
integrated network
mixed on-off policy exploration
sample complexity

Journal

IEEE Transactions on Neural Networks and Learning Systems cover
IEEE Transactions on Neural Networks and Learning Systems
IF:
8.9
Papers:
7.5K
Citations:
7.2W

Organization

S
shenzhen institute of advanced technology, cas
Scholars:
5.6K
Papers: 4.5K
Citations: 7
C
chinese academy of sciences
Scholars:
56.3W
Papers: 44.8W
Citations: 704