Return
Efficient Classification-Based Constraints for Offline Reinforcement Learning
DOI:10.3390/app152212197.png)
Abstract
En 中文
Existing distribution-based constraints in offline reinforcement learning, such as bootstrapping error accumulation reduction (BEAR), can prevent policy deviation but often ignore action quality, leading to O(n2) complexity and limited interpretability. This study introduces a classification-based approach that employs pairwise action quality comparison to replace complex distributional constraints. A binary classifier learns the relative quality of two actions in the same state by comparing their Q-values, prioritizing value-aware selection while maintaining conservative behavior. Experiments on benchmark environments demonstrate that the proposed method consistently improves upon BEAR, achieving 3x on average and up to 5x in some environments. The algorithm reduces computational complexity from O(n2) to O(n) while providing intuitive monitoring through classifier accuracy. These results indicate that efficient quality-based comparisons can serve as a practical and efficient alternative to computationally expensive distributional constraints in offline reinforcement learning, yielding practical gains in both performance and scalability.
Keywords:
offline reinforcement learning
classification-based constraints
action quality comparison
conservative policy optimization
computational efficiency
Q-value estimation
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
A
IF:
2.5
Papers:
7.6K
Citations:
4
Organization
Cited Papers
Classification-Based Q-Value Estimation for Continuous Actor-Critic Reinforcement Learning
Symmetry
IF0
no more


