arrow
返回

Beyond backpropagate through time: Efficient model-based training through time-splitting

delete2022-05-24
delete0
PRE
AI
J
Jiaxin Gao
Y
Yang Guan
李温玉 封面图
李温玉 (Wenyu Li)
S
Shengbo Eben Li *
F
Fei Ma
J
Jianfeng Zheng
J
Junqing Wei
张博 封面图
张博 (Bo Zhang)
K
Keqiang Li
DOI:10.1002/int.22928delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Model-based policy gradient (MBPG) has been employed to seek an approximate solution to the optimal control problem. However, there is coupling between adjacent states due to temporal dependencies, making the training time grow linearly with the time horizon. This paper reshapes the training process of MBPG with the time-splitting technique to establish a time-independent algorithm called Training Through Time-Splitting (T3S). First, copy the coupled variables to obtain two independent variables. Meanwhile, an extra variable together with an equivalence constraint is introduced for problem consistency. Then, the transformed problem divides into subproblems with carefully derived loss functions. Subproblems own decoupled variables and shared policy networks, which means they can be optimized concurrently. Guided by the algorithm design, this paper further proposes an asynchronous parallel training scheme to accelerate training efficiency. Numerical simulation shows that the T3S algorithm outperforms the MBPG algorithm by 83.6% in wall-clock time with a trajectory tracking task.
Keyword:
model-based policy gradient
optimal control
parallel training
reinforcement learning
time-splitting

期刊

International Journal of Intelligent Systems 封面图
International Journal of Intelligent Systems
IF:
3.7
论文数:
3.0K
被引数:
8.1K

机构

T
tsinghua university
学者数:
11.9W
论文数: 10.0W
被引数: 137
N
nankai university
学者数:
4.8W
论文数: 3.3W
被引数: 74