arrow
Return

TS-BEV: BEV object detection algorithm based on temporal-spatial feature fusion

delete2024-09-01
delete8
PRE
AI
X
Xinlong Dong
P
Peicheng Shi *
H
Heng Qi
A
Aixi Yang
T
Taonian Liang
DOI:10.1016/j.displa.2024.102814delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In order to accurately identify occluding targets and infer the motion state of objects, we propose a Bird's-Eye View Object Detection Network based on Temporal-Spatial feature fusion (TS-BEV), which replaces the previous multi-frame sampling method by using the cyclic propagation mode of historical frame instance information. We design a new Temporal-Spatial feature fusion attention module, which fully integrates temporal information and spatial features, and improves the inference and training speed. In response to realize multi-frame feature fusion across multiple scales and views, we propose an efficient Temporal-Spatial deformable aggregation module, which performs feature sampling and weighted summation from multiple feature maps of historical frames and current frames, and makes full use of the parallel computing capabilities of GPUs and AI chips to further improve efficiency. Furthermore, in order to solve the lack of global inference in the context of temporal-spatial fusion BEV features and the inability of instance features distributed in different locations to fully interact, we further design the BEV self-attention mechanism module to perform global operation of features, enhance global inference ability and fully interact with instance features. We have carried out extensive experimental experiments on the challenging BEV object detection nuScenes dataset, quantitative results show that our method achieves excellent performance of 61.5% mAP and 68.5% NDS in camera-only 3D object detection tasks, and qualitative results show that TS-BEV can effectively solve the problem of 3D object detection in complex traffic background with lack of light at night, with good robustness and scalability.
Keywords:
BEV object detection
Temporal-Spatial feature fusion
Temporal-Spatial deformable aggregation
BEV self-attention

Journal

Displays cover
Displays
IF:
3.4
Papers:
2.1K
Citations:
3.2K

Organization

A
Anhui Polytechnic University
Scholars:
3.8K
Papers: 2.5K
Citations: 3.5K
Z
zhejiang university
Scholars:
17.5W
Papers: 12.0W
Citations: 152
W
wuhan university
Scholars:
8.0W
Papers: 5.8W
Citations: 70
researcher View more organizations