arrow
Return

Reinforcement Learning Decision and Planning Algorithm Using Posterior Return Heuristic

delete2025-09-26
delete0
PRE
AI
Z
Zhengtang Ma
E
Eksan Firkat
X
Xiaming Yuan
Y
Yijian Duan
J
Jihong Zhu
A
Askar Hamdulla
DOI:10.1109/TVT.2025.3614880delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The complexities and dynamics of modern driving environments have amplified the uncertainties faced by autonomous driving systems, posing significant challenges to their decision-making processes. This paper introduces a novel reinforcement learning framework inspired by posterior return estimation, designed to enhance safety and foresight in robotic decision-making. To overcome the challenges of long-term dependency in reinforcement learning, a tree structure is utilized to record the interaction history between the agent and the environment. The posterior distribution of returns for each state is estimated from this historical data, with a neural network employed to model this distribution for enhanced decision-making. The framework's efficacy is validated through two decision-making and planning scenarios, where it surpasses several commonly used frameworks and demonstrates notably more anticipatory decision-making patterns.
Keywords:
Reinforcement learning
autonomous driving
decision making
posterior estimation

Journal

IEEE Transactions on Vehicular Technology cover
IEEE Transactions on Vehicular Technology
IF:
7.1
Papers:
1.8W
Citations:
6.6W

Organization

G
Guangxi University
Scholars:
4.1K
Papers: 1.3K
Citations: 3.2W
T
tsinghua university
Scholars:
11.8W
Papers: 10.0W
Citations: 137
T
Tsinghua University
Scholars:
8.6K
Papers: 4.1K
Citations: 17.7W
X
Xinjiang University
Scholars:
1.4W
Papers: 8.7K
Citations: 1.1W
researcher View more organizations