Return
Reinforcement Learning Decision and Planning Algorithm Using Posterior Return Heuristic
DOI:10.1109/TVT.2025.3614880.png)
Abstract
En 中文
The complexities and dynamics of modern driving environments have amplified the uncertainties faced by autonomous driving systems, posing significant challenges to their decision-making processes. This paper introduces a novel reinforcement learning framework inspired by posterior return estimation, designed to enhance safety and foresight in robotic decision-making. To overcome the challenges of long-term dependency in reinforcement learning, a tree structure is utilized to record the interaction history between the agent and the environment. The posterior distribution of returns for each state is estimated from this historical data, with a neural network employed to model this distribution for enhanced decision-making. The framework's efficacy is validated through two decision-making and planning scenarios, where it surpasses several commonly used frameworks and demonstrates notably more anticipatory decision-making patterns.
Keywords:
Reinforcement learning
autonomous driving
decision making
posterior estimation
Journal
IF:
7.1
Papers:
1.8W
Citations:
6.6W

