arrow
Return

Candidate ratio guided proximal policy optimization

delete2025-07-01
delete0
PRE
AI
Y
Yao Li
谭晓阳 (Xiaoyang Tan) *
DOI:10.1016/j.engappai.2025.110576delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Proximal Policy Optimization (PPO) is a well-studied policy search method that improves policies monotonically by enforcing the probability ratio to remain within a clipping range. The probability ratios used to measure the policy distance change dynamically during policy optimization. However, the clipping range remains fixed throughout the policy training stage. A fixed clipping range may lead to sample inefficiency or aggressive policy updates, negatively affecting policy performance. In this paper, we propose a candidate-ratio-guided Proximal Policy Optimization method with self-adaptive clipping ratios to design the clipping range for improved sample efficiency and monotonic policy improvement. The clipping ratio is adjusted based on the average candidate ratios derived from actions sampled around the policy suggested actions. Increasing the clipping ratio allows more collected data to be used for policy optimization, while decreasing it effectively enforces probability ratios within the clipping range for monotonic policy improvement. Experimental results on MuJoCo tasks show that our method achieves more stable performance compared to existing baselines.
Keywords:
Reinforcement learning
Policy gradient
Conservative policy iteration
Proximal policy optimization

Journal

Engineering Applications of Artificial Intelligence cover
Engineering Applications of Artificial Intelligence
IF:
8
Papers:
5.4K
Citations:
3.5W

Organization

N
Nanjing University of Aeronautics and Astronautics
Scholars:
7.4K
Papers: 3.1K
Citations: 2.4W