arrow
Return

Temporal and spatial context aware voxel transformer for semantic scene completion

delete2026-02-28
delete0
PRE
AI
吴旖琦 cover
吴旖琦 (Yiqi Wu)
C
Changliang Li
J
Jiale He
F
Feng Cao
C
Cuilian Lei
Y
Yilin Chen
张德军 (Dejun Zhang)
DOI:10.1016/j.neunet.2026.108754delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Semantic Scene Completion (SSC) aims to recover comprehensive 3D structure and semantic understanding from partial visual observations, playing a crucial role in autonomous driving perception. However, current camera-based SSC methods often suffer from insufficient temporal reasoning and unreliable depth estimation, which results in incomplete geometry recovery and ambiguous semantic interpretation. To address these challenges, we develop a method that recovers complete 3D semantics by progressively integrating multi-frame context and depth cues. Starting from input images, contextual and depth-related features are separately extracted. Contextual features are then temporally aligned through temporal-spatial aware mechanisms that maintain coherence under viewpoint variations. Meanwhile, monocular depth priors are probabilistically fused with stereo-derived constraints to refine depth estimates. The refined depth further guides volumetric reconstruction, enabling accurate and stable semantic scene completion. The proposed method is evaluated on SemanticKITTI and SSCBench-KITTI-360, achieving results on par with or surpassing current state-of-the-art methods. The source code is available at: https://www.github.com/djzgroup/SSC .
Keywords:
Semantic Scene Completion
Temporal Context
Depth Estimation
Voxel Transformer
Semantic Understanding

Journal

Neural Networks cover
Neural Networks
IF:
6.3
Papers:
7.8K
Citations:
3.0W

Organization

C
china university of geosciences
Scholars:
8.1K
Papers: 3.0K
Citations: 0
W
wuhan institute of technology
Scholars:
1.0W
Papers: 6.5K
Citations: 11