arrow
返回

Solving time-delay issues in reinforcement learning via transformers

delete2024-09-10
delete0
PRE
AI
B
Bo Xia
Z
Zaihui Yang
M
Minzhi Xie
Y
Yongzhe Chang
B
Bo Yuan
Z
Zhiheng Li
王学谦 (Xueqian Wang) *
B
Bin Liang
DOI:10.1007/s10489-024-05830-2delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The presence of observation and action delays in remote control scenarios significantly challenges the decision-making of agents that depend on immediate interactions, particularly within traditional deep reinforcement learning (DRL) algorithms. Existing approaches attempt to tackle this problem through various strategies, such as predicting delayed states, transforming delayed Markov Decision Processes (MDPs) into delay-free equivalents. However, both model-free and model-based methods require extensive online data, making them time-consuming and resource-intensive. To effectively handle time-delay challenges and develop a competent and robust RL algorithm, the Augmented Decision Transformer (ADT) is proposed as the first offline RL algorithm designed to enable agents to manage diverse tasks with various constant delays. It transforms a deterministic delayed MDP (DDMDP) into a standard MDP by simulating trajectories in delayed environments using offline dataset from undelayed environments. The Decision Transformer, an autoregressive model, is then employed to train a decision model based on expected rewards, past state sequences and past action sequences. Extensive experiments conducted on MuJoCo and Adroit tasks validate the robustness and efficiency of the ADT, with its average performance across all tasks being 56% better than the worst-performing comparative algorithms. The results demonstrate that the ADT can outperform state-of-the-art RL counterparts, achieving superior performance across various tasks with different delay conditions.
Keyword:
Deep reinforcement learning
Time delay
Deterministic delayed Markov Decision Process
Offline reinforcement learning
Decision transformer

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

T
tsinghua university
学者数:
11.9W
论文数: 10.0W
被引数: 137
T
Tsinghua Shenzhen International Graduate School
学者数:
6.8K
论文数: 4.9K
被引数: 9
引用论文

引用论文

err分享
err收藏
err分享
err收藏
Delay-aware model-based reinforcement learning for continuous control
err2021-08-01
err36
errOAAI
errChen, Baiming; Xu, Mengdi; Li, Liang; Zhao, Ding
err分享
err收藏
学者 查看更多内容