arrow
Return

State transition difference prediction for deep reinforcement learning

delete2025-12-04
delete0
PRE
AI
H
Haotian Chi
刘兆赓 cover
刘兆赓 (Zhaogeng Liu)
X
Xing Chen
B
Bohao Qu
J
Jifeng Hu
Y
Yuan Jiang
陈贺昌 cover
陈贺昌 (Hechang Chen)
Y
Yi Chang
DOI:10.1016/j.patcog.2025.112824delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• We propose STDP, a novel RL method that learns latent dynamics by predicting transition differences between states. • STDP-E and STDP-I model explicit and implicit differences in state transitions, enabling complementary learning strategies. • STDP achieves superior sample efficiency over baselines on MuJoCo and D4RL in both online and offline RL tasks. • In-depth ablation and hyper-parameter analysis verify the effectiveness and robustness of STDP-E and STDP-I.

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

N
Nanyang Technological University
Scholars:
4.9W
Papers: 4.8W
Citations: 8.1W
J
Jilin University
Scholars:
8.7W
Papers: 5.5W
Citations: 8.9K