arrow
Return

Dyna-style Model-based reinforcement learning with Model-Free Policy Optimization

delete2024-03-01
delete0
PRE
AI
董坤 cover
董坤 (Kun Dong)
Y
Yongle Luo
Y
Yuxin Wang
Y
Yu Liu
C
Chengeng Qu
Q
Q. Zhang
E
Erkang Cheng
Z
Zhiyong Sun
宋博 cover
宋博 (Bo Song) *
DOI:10.1016/j.knosys.2024.111428delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Dyna-style Model-based reinforcement learning (MBRL) methods have demonstrated superior sample efficiency compared to their model-free counterparts, largely attributable to the leverage of learned models. Despite these advancements, the effective application of these learned models remains challenging, largely due to the intricate interdependence between model learning and policy optimization, which presents a significant theoretical gap in this field. This paper bridges this gap by providing a comprehensive theoretical analysis of Dyna-style MBRL for the first time and establishing a return bound in deterministic environments. Building upon this analysis, we propose a novel schema called Model-Based Reinforcement Learning with Model-Free Policy Optimization (MBMFPO). Compared to existing MBRL methods, the proposed schema integrates modelfree policy optimization into the MBRL framework, along with some additional techniques. Experimental results on various continuous control tasks demonstrate that MBMFPO can significantly enhance sample efficiency and final performance compared to baseline methods. Furthermore, extensive ablation studies provide robust evidence for the effectiveness of each individual component within the MBMFPO schema. This work advances both the theoretical analysis and practical application of Dyna-style MBRL, paving the way for more efficient reinforcement learning methods.
Keywords:
Reinforcement learning
Robotics
Data efficiency

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

C
chinese academy of sciences
Scholars:
56.5W
Papers: 44.9W
Citations: 704