Return
Offline Generalized Actor-Critic with Distance Regularization
DOI:10.1109/JAS.2025.125633.png)
Abstract
En 中文
In order to address the issue of overly conservative offline reinforcement learning (RL) methods that limit the generalization of policy in the out-of-distribution (OOD) region, this article designs a surrogate target for OOD value function based on dataset distance and proposes a novel generalized Q-learning mechanism with distance regularization (GQDR). In theory, we not only prove the convergence of GQDR, but also ensure that the difference between the Q-value learned by GQDR and its true value is bounded. Furthermore, an offline generalized actor-critic method with distance regularization (OGACDR) is proposed by combining GQDR with actor-critic learning framework. Two implementations of OGACDR, OGACDR-EXP and OGACDR-SQR, are introduced according to exponential (EXP) and open-square (SQR) distance weight functions, and it has been theoretically proved that OGACDR provides a safe policy improvement. Experimental results on Gym-MuJoCo continuous control tasks show that OGACDR can not only alleviate the overestimation and overconservatism of Q-value function, but also outperform conservative offline RL baselines.
Keywords:
Actor-critic
distance regularization
generalized Q-learning
offline reinforcement learning
out-of-distribution (ODD)
Journal
I
IF:
0
Papers:
116
Citations:
0

