arrow
Return

Uncertainty-based bootstrapped optimization for offline reinforcement learning

delete2024-11-02
delete0
PRE
AI
T
Tianyi Li
G
Genke Yang *
J
Jian Chu
DOI:10.1007/s13042-024-02439-2delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Offline reinforcement learning (offline RL) promises to learn effective policies from previously-collected, static datasets without offering further possibility for exploration. However, offline RL encounters significant challenges primarily due to algorithmic difficulties arising from function approximation errors caused by extrapolating from out-of-distribution (OOD) data points. In this work, we propose uncertainty-based bootstrapped optimization (UBO), which aims to address the distributional shift induced by the fixed datasets. First, we take advantage of the bootstrapped architecture to implicitly approximate the epistemic uncertainty for the training instances. Then, we apply both the implicit and explicit penalties to the OOD data with high prediction uncertainties. Finally, we introduce a training paradigm based on the upper confidence bound (UCB) strategy for the bootstrapping updates, which enables the algorithm to thoroughly assess the varying performance of each bootstrapped head. We compare UBO with other prevailing offline RL algorithms on D4RL benchmarks. Experiments on various tasks demonstrate that the proposed algorithm can outperform or be competitive with the previous state-of-the-art on most of the tasks.
Keywords:
Offline RL
Bootstrapped learning
Distributional shift
Uncertainty optimization

Journal

International Journal of Machine Learning and Cybernetics cover
International Journal of Machine Learning and Cybernetics
IF:
2.7
Papers:
3.1K
Citations:
5.6K

Organization

S
shanghai jiao tong university
Scholars:
15.6W
Papers: 11.6W
Citations: 159