arrow
Return

Offline spatial-temporal actor-critic learning for robot crowd navigation without behavior regularization

delete2025-12-19
delete0
PRE
AI
S
Shuai Zhou
符浩 (Hao Fu) *
H
He, Haodong
W
Wei Liu
Z
Zixin Huang
DOI:10.1007/s11370-025-00664-4delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Robot crowd navigation has been gaining increasing attention and popularity in various practical applications. In the existing research, deep reinforcement learning has been applied to robot crowd navigation by training policies in an online mode. However, this inevitably leads to unsafe exploration and consequently causes low sampling efficiency during pedestrian-robot interaction. To this end, this paper proposes an offline spatial-temporal actor-critic algorithm for robot crowd navigation by utilizing pre-collected crowd navigation experience. Specifically, a novel SARSA-like loss function is defined to update the separate value function, eliminating the need for behavior regularization or constraints. Furthermore, this algorithm incorporates a spatial-temporal transformer into offline actor-critic learning that captures the spatial-temporal features from the offline pedestrian-robot interactions. It allows our robot navigation policy to exhibit greater adaptability and less conservatism to the highly dynamic crowd environments. Experimental results demonstrate that the proposed algorithm outperforms the state-of-the-art methods by means of qualitative and quantitative analysis.
Keywords:
Robot crowd navigation
Deep reinforcement learning
Offline learning
Neural networks

Journal

Intelligent Service Robotics cover
Intelligent Service Robotics
IF:
4.3
Papers:
78
Citations:
1.2K

Organization

W
wuhan institute of technology
Scholars:
1.0W
Papers: 6.5K
Citations: 11