Return
Candidate ratio guided proximal policy optimization
DOI:10.1016/j.engappai.2025.110576.png)
Abstract
En 中文
Proximal Policy Optimization (PPO) is a well-studied policy search method that improves policies monotonically by enforcing the probability ratio to remain within a clipping range. The probability ratios used to measure the policy distance change dynamically during policy optimization. However, the clipping range remains fixed throughout the policy training stage. A fixed clipping range may lead to sample inefficiency or aggressive policy updates, negatively affecting policy performance. In this paper, we propose a candidate-ratio-guided Proximal Policy Optimization method with self-adaptive clipping ratios to design the clipping range for improved sample efficiency and monotonic policy improvement. The clipping ratio is adjusted based on the average candidate ratios derived from actions sampled around the policy suggested actions. Increasing the clipping ratio allows more collected data to be used for policy optimization, while decreasing it effectively enforces probability ratios within the clipping range for monotonic policy improvement. Experimental results on MuJoCo tasks show that our method achieves more stable performance compared to existing baselines.
Keywords:
Reinforcement learning
Policy gradient
Conservative policy iteration
Proximal policy optimization
Journal
IF:
8
Papers:
5.4K
Citations:
3.5W

