Return
State transition difference prediction for deep reinforcement learning
DOI:10.1016/j.patcog.2025.112824.png)
Abstract
En 中文
• We propose STDP, a novel RL method that learns latent dynamics by predicting transition differences between states. • STDP-E and STDP-I model explicit and implicit differences in state transitions, enabling complementary learning strategies. • STDP achieves superior sample efficiency over baselines on MuJoCo and D4RL in both online and offline RL tasks. • In-depth ablation and hyper-parameter analysis verify the effectiveness and robustness of STDP-E and STDP-I.
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W

