Return
Reward shaping using graph MAMBA
J
G
Y
杨
J
DOI:10.1016/j.patcog.2026.114404.png)
Abstract
En 中文
• Propose a temporal graph MAMBA architecture for reward shaping without changing the optimal policy. • Develop a kernelized message-passing mechanism that reduces the computational complexity from quadratic to linear time. • Attain state-of-the-art performance across multiple standard benchmark platforms.
Keywords:
Markov decision process
Reinforcement learning
Reward shaping
Graph
MAMBA
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W
