arrow
返回

Enhancing model learning in reinforcement learning through Q-function-guided trajectory alignment

delete2025-05-01
delete0
PRE
AI
X
Xin Du
S
Shan Zhong
S
Shengrong Gong *
Z
Zhenyu Qi
DOI:10.1007/s10489-024-06083-9delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Model-based reinforcement learning (MBRL) methods hold great promise for achieving excellent sample efficiency by fitting a dynamics model to previously observed data and leveraging it for RL or planning. However, the resulting trajectories may diverge from actual-world trajectories due to the accumulation of errors in multi-step model sampling, particularly for longer horizons. This undermines the performance of MBRL and significantly affects sample efficiency. Therefore, we present a trajectory alignment capable of aligning simulated trajectories with their real counterparts from any initial random state and with adaptive length, enabling the preparation of paired real-simulated samples to minimize compounding errors. Additionally, we design a Q-function function to estimate Q values for the paired real-simulated samples. The simulated samples whose Q-value difference from the real ones surpasses a given threshold will be discarded, thus preventing the model from over-fitting to erroneous samples. Experimental results demonstrate that both trajectory alignment and Q-function guided sample filtration contribute to improving policy and sample efficiency. Our method surpasses previous state-of-the-art model-based approaches in both sample efficiency and asymptotic performance across a series of challenging control tasks. The code is open source and available at https://github.com/duxin0618/qgtambpo.git.
Keyword:
Reinforcement learning
Dynamics model
Model-based RL algorithms
Model learning

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

C
Changshu Institute of Technology
学者数:
2.2K
论文数: 1.8K
被引数: 3
U
University of Arizona
学者数:
3.6W
论文数: 3.2W
被引数: 980
引用论文

引用论文

暂无论文信息