arrow
Return

PACER: A fully push-forward-based distributional reinforcement learning algorithm

delete2025-08-19
delete0
delete
OA
AI
W
Wensong Bai
C
Chao Zhang *
Y
Yichao Fu
P
Peilin Zhao
H
Hui Qian
DOI:10.1016/j.neucom.2025.131301delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
In this paper, we propose the first fully push-forward-based distributional reinforcement learning algorithm, named PACER, which consists of a distributional critic, a stochastic actor and a sample-based encourager. Specifically, the push-forward operator is leveraged in both the critic and actor to model the return distributions and stochastic policies respectively, enabling them with equal modeling capability and thus enhancing the synergetic performance. Since it is infeasible to obtain the density function of the push-forward policies, novel sample-based regularizers are integrated into the encourager to incentivize efficient exploration and alleviate the risk of trapping into local optima. Moreover, a sample-based stochastic utility value policy gradient is established for the push-forward policy update, which circumvents the explicit demand for the policy density function in existing REINFORCE-based stochastic policy gradients. As a result, PACER fully utilizes the modeling capability of the push-forward operator and is able to explore a broader class of the policy spaces compared with limited policy classes used in existing distributional actor-critic algorithms (i.e. Gaussians). We validate the critical role of each component in our algorithm with extensive empirical studies. Experimental results demonstrate the superiority of our algorithm over state-of-the-art methods, achieving an average score improvement of 10 % on the Mujoco continuous control benchmark.
Keywords:
Distributional reinforcement learning
Actor-critic
Push-forward policy
Sample-based regularizer
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

C
College of Computer Science and Technology
Scholars:
850
Papers: 296
Citations: 0
S
School of Artificial Intelligence
Scholars:
754
Papers: 344
Citations: 0