arrow
返回

Generative Adversarial Soft Actor-Critic

delete2024-01-01
delete0
PRE
AI
H
Hyo-Seok Hwang
Y
Yoojoong Kim *
J
Junhee Seok *
DOI:10.1109/TNNLS.2024.3493113delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In deep reinforcement learning (RL), learning stochastic and multimodal policies are crucial for various tasks, but most continuous control algorithms model the policy using a deterministic or unimodal Gaussian distribution. Despite being designed to inherently learn a stochastic policy under the maximum entropy RL framework, soft actor-critic (SAC) also uses a factorized Gaussian policy for tractable optimization, which not only restricts the expressiveness of the policy but also ignores correlations among the components of the action vector. In this article, we revisit the approach of employing normalizing flow for SAC policy, justified by the change of variable theorem, and then propose a state-dependent nonvolume preserving (SD-NVP) architecture suitable for SAC learning. In addition, we introduce a generative adversarial SAC (GASAC) that implicitly defines and optimizes various divergences without calculating the normalization constant through a generative adversarial loss. Experimental results on multigoal environment and MuJoCo continuous control tasks suite demonstrate that GASAC model multimodal policy and learn policy more stably in terms of cumulative return.
Keyword:
Entropy
Mathematical models
Generative adversarial networks
Training
Stochastic processes
Vectors
Optimization
Measurement
Computational modeling
Predictive models
Continuous control
deep learning
deep reinforcement learning (RL)
generative model

期刊

IEEE Transactions on Neural Networks and Learning Systems 封面图
IEEE Transactions on Neural Networks and Learning Systems
IF:
8.9
论文数:
7.6K
被引数:
7.2W

机构

K
Korea University
学者数:
3.6W
论文数: 3.8W
被引数: 4.4W
C
catholic university of korea
学者数:
1.4W
论文数: 1.3W
被引数: 9
引用论文

引用论文

暂无论文信息