arrow
Return

LiDAR video object segmentation with dynamic kernel refinement

delete2024-02-01
delete1
PRE
AI
J
Jianbiao Mei
Y
Yu Yang
M
Mengmeng Wang
Z
Zizhang Li
J
J.B. Ra
刘勇 (Yong Liu) *
DOI:10.1016/j.patrec.2023.12.013delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In this paper, we formalize memory-and tracking-based methods to perform the LiDAR-based Video Object Segmentation (VOS) task, which segments points of the specific 3D target (given in the first frame) in a LiDAR sequence. LiDAR-based VOS can directly provide target-aware geometric information for practical application scenarios like behavior analysis and anticipating danger. We first construct a LiDAR-based VOS dataset named KITTI-VOS based on SemanticKITTI, which acts as a testbed and facilitates comprehensive evaluations of algorithm performance. Next, we provide two types of baselines, i.e., memory-based and tracking-based baselines, to explore this task. Specifically, the first memory-based pipeline is built on a space-time memory network equipped with the non-local spatiotemporal attention-based memory bank. We further design a more potent variant to introduce the locality into the spatiotemporal attention module by local self-attention and cross-attention modules. For the second tracking-based baseline, we modify two representative 3D object tracking methods to adapt to LiDAR-based VOS tasks. Finally, we propose a refine module that takes mask priors and generates object-aware kernels, which could boost all the baselines' performance. We evaluate the proposed methods on the dataset and demonstrate their effectiveness.
Keywords:
LiDAR segmentation
Video object segmentation
Dynamic kernel

Journal

Pattern Recognition Letters cover
Pattern Recognition Letters
IF:
3.3
Papers:
7.8K
Citations:
1.6W

Organization

Z
zhejiang university
Scholars:
17.5W
Papers: 12.0W
Citations: 152