Return
PACER: A fully push-forward-based distributional reinforcement learning algorithm
DOI:10.1016/j.neucom.2025.131301.png)
Abstract
En 中文
In this paper, we propose the first fully push-forward-based distributional reinforcement learning algorithm, named PACER, which consists of a distributional critic, a stochastic actor and a sample-based encourager. Specifically, the push-forward operator is leveraged in both the critic and actor to model the return distributions and stochastic policies respectively, enabling them with equal modeling capability and thus enhancing the synergetic performance. Since it is infeasible to obtain the density function of the push-forward policies, novel sample-based regularizers are integrated into the encourager to incentivize efficient exploration and alleviate the risk of trapping into local optima. Moreover, a sample-based stochastic utility value policy gradient is established for the push-forward policy update, which circumvents the explicit demand for the policy density function in existing REINFORCE-based stochastic policy gradients. As a result, PACER fully utilizes the modeling capability of the push-forward operator and is able to explore a broader class of the policy spaces compared with limited policy classes used in existing distributional actor-critic algorithms (i.e. Gaussians). We validate the critical role of each component in our algorithm with extensive empirical studies. Experimental results demonstrate the superiority of our algorithm over state-of-the-art methods, achieving an average score improvement of 10 % on the Mujoco continuous control benchmark.
Keywords:
Distributional reinforcement learning
Actor-critic
Push-forward policy
Sample-based regularizer
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

