arrow
Return

Efficient Classification-Based Constraints for Offline Reinforcement Learning

delete2025-11-17
delete0
delete
OA
AI
C
Chayoung Kim *
DOI:10.3390/app152212197delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Existing distribution-based constraints in offline reinforcement learning, such as bootstrapping error accumulation reduction (BEAR), can prevent policy deviation but often ignore action quality, leading to O(n2) complexity and limited interpretability. This study introduces a classification-based approach that employs pairwise action quality comparison to replace complex distributional constraints. A binary classifier learns the relative quality of two actions in the same state by comparing their Q-values, prioritizing value-aware selection while maintaining conservative behavior. Experiments on benchmark environments demonstrate that the proposed method consistently improves upon BEAR, achieving 3x on average and up to 5x in some environments. The algorithm reduces computational complexity from O(n2) to O(n) while providing intuitive monitoring through classifier accuracy. These results indicate that efficient quality-based comparisons can serve as a practical and efficient alternative to computationally expensive distributional constraints in offline reinforcement learning, yielding practical gains in both performance and scalability.
Keywords:
offline reinforcement learning
classification-based constraints
action quality comparison
conservative policy optimization
computational efficiency
Q-value estimation
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

A
Applied Sciences-Basel
IF:
2.5
Papers:
7.6K
Citations:
4

Organization

Hankyong National University cover
Hankyong National University
Scholars:
840
Papers: 1.0K
Citations: 759
Cited Papers

Cited Papers

Human-level control through deep reinforcement learning
err2015-02-25
err0
PREAI
errVolodymyr Mnih; Koray Kavukcuoglu; David Silver; Andrei A. Rusu; Joel Veness; Marc G. Bellemare; Alex Graves; Martin Riedmiller; Andreas K. Fidjeland; Georg Ostrovski; Stig Petersen; Charles Beattie; Amir Sadik; Ioannis Antonoglou; Helen King; Dharshan Kumaran; Daan Wierstra; Shane Legg; Demis Hassabis
errShare
errSave
err1999-01-01
err0
PREAI
errPedro Domingos
errShare
errSave
Learning to rank
err2007-06-20
err0
PREAI
errZhe Cao; Tao Qin; Tie-Yan Liu; Ming-Feng Tsai; Hang Li
errShare
errSave
Learning to rank using gradient descent
err2005-01-01
err0
PREAI
errChris Burges; Tal Shaked; Erin Renshaw; Ari Lazier; Matt Deeds; Nicole Hamilton; Greg Hullender
errShare
errSave
Batch Reinforcement Learning
err2012-01-01
err0
PREAI
errSascha Lange; Thomas Gabel; Martin Riedmiller
errShare
errSave
no more