arrow
Return

Dual Behavior Regularized Offline Deterministic Actor-Critic

delete2024-08-01
delete0
PRE
AI
S
Shuo Cao
X
Xuesong Wang
Y
Yuhu Cheng *
DOI:10.1109/TSMC.2024.3388007delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
To mitigate the extrapolation error arising from offline reinforcement learning (RL) paradigm, existing methods typically make learned Q-functions over-conservative or enforce global policy constraints. In this article, we propose a dual behavior regularized offline deterministic Actor-Critic (DBRAC) by simultaneously performing behavior regularization on the coupling-iterative policy evaluation (PE) and policy improvement (PI) in the policy iteration process. In the PE phase, the difference between the Q-function and behavior value is first taken as the anti-exploration behavior value regularization term to drive the Q-function toward its true Q-value, which significantly reduces the conservatism of learned Q-function. In the PI phase, the estimated action variances of behavior policy in different states are then utilized for designing the weight and threshold of mild-local behavior cloning regularization term, which standardizes the local improvement potential of learned policy. Experiments on the well-known datasets for deep data-driven RL (D4RL) demonstrate that the DBRAC can quickly learn more competitive task-solving policies in various offline situations with different data qualities, significantly outperforming state-of-the-art offline RL baselines.
Keywords:
Anti-exploration behavior value
dual behavior regularization (DBR)
mild-local behavior cloning (BC)
offline deterministic Actor-Critic
reinforcement learning (RL)

Journal

IEEE Transactions on Cybernetics cover
IEEE Transactions on Cybernetics
IF:
10.5
Papers:
1.1W
Citations:
5.0W

Organization

No organization information available