arrow
Return

SF3D-MDA: A Road-User Behavior Annotation Framework and Recognition Method Based on Vehicle-Mounted Visual Information

delete2025-07-07
delete0
PRE
AI
C
C. X. Yue
L
Lisheng Jin
B
Baicang Guo
张宏玉 (Hongyu Zhang)
J
Junchen Liu
X
X.H. Liu
C
Chuanqiang An
DOI:10.1109/JSEN.2025.3584066delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Addressing the critical challenge of spatiotemporal semantic disjunction caused by conventional bird’s-eye view (BEV) trajectory modeling methods in ego-vehicle perspective road-user behavior recognition tasks, this article proposes a dedicated ego-vehicle-centric road-user behavior recognition methodology as a systematic solution. The methodology introduces the behavior–location joint annotation (BLJA) framework to address the inadequacy of the existing annotation framework in accommodating multimodal interaction dynamics and the scarcity of dedicated ego-vehicle perspective road-user behavior datasets. Furthermore, we construct our ROAD-Waymo-trans dataset using this framework, semantically enriched with fine-grained behavioral and positional labels to align with autonomous driving perception requirements. We propose a SlowFast-based 3-D bidirectional feature pyramid network (FPN) with multidimensional attention (SF3D-MDA) mechanisms. The model employs a dual-path 3-D bidirectional FPN (3-DBi-FPN): the slow pathway embeds channel–spatial attention (CSA) to capture pixel-level spatial details, while the fast pathway integrates channel–temporal attention (CTA) to model cross-frame spatiotemporal dependencies. Experimental results on the ROAD dataset show frame-level f-mAP@0.5 improvements of 1.7% (agent), 4% (action), and 3% (location). On the ROAD-Waymo dataset, video-level v-mAP@0.2 achieves gains of 7.4% (agent), 1.8% (action), and 2.7% (location). Ablation studies confirm that multidimensional attention mechanisms contribute 3.59%–4.37% performance boosts across tasks. Experimental results demonstrate that the improved model exhibits robustness in modeling multiagent interactions within complex traffic scenarios and produces accurate detection results. The proposed method successfully bridges the gap between perception and decision-making through context-aware annotations and spatiotemporally coherent architecture design. The results validate its utility for reliable autonomous driving systems in real-world environments. The executable codebase central to this research is available at the following repository: <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/Hilbert-space007/SF3D-MDA</uri>
Keywords:
Attention mechanism
autonomous driving
dual-path 3-D network
road-user behavior recognition
sensor applications (APPLs)
spatiotemporal dependency feature fusion

Journal

IEEE Sensors Journal cover
IEEE Sensors Journal
IF:
4.5
Papers:
2.1W
Citations:
7.3W

Organization

Y
yanshan university
Scholars:
4.1K
Papers: 1.3K
Citations: 0