Return
Temporal and spatial context aware voxel transformer for semantic scene completion
DOI:10.1016/j.neunet.2026.108754.png)
Abstract
En 中文
Semantic Scene Completion (SSC) aims to recover comprehensive 3D structure and semantic understanding from partial visual observations, playing a crucial role in autonomous driving perception. However, current camera-based SSC methods often suffer from insufficient temporal reasoning and unreliable depth estimation, which results in incomplete geometry recovery and ambiguous semantic interpretation. To address these challenges, we develop a method that recovers complete 3D semantics by progressively integrating multi-frame context and depth cues. Starting from input images, contextual and depth-related features are separately extracted. Contextual features are then temporally aligned through temporal-spatial aware mechanisms that maintain coherence under viewpoint variations. Meanwhile, monocular depth priors are probabilistically fused with stereo-derived constraints to refine depth estimates. The refined depth further guides volumetric reconstruction, enabling accurate and stable semantic scene completion. The proposed method is evaluated on SemanticKITTI and SSCBench-KITTI-360, achieving results on par with or surpassing current state-of-the-art methods. The source code is available at: https://www.github.com/djzgroup/SSC .
Keywords:
Semantic Scene Completion
Temporal Context
Depth Estimation
Voxel Transformer
Semantic Understanding
Journal
IF:
6.3
Papers:
7.8K
Citations:
3.0W

