Return
Semantic Decoupling Based Semantic Scene Completion From a Single Depth Image
DOI:10.1109/LRA.2025.3634907.png)
Abstract
En 中文
Semantic Scene Completion (SSC) is a task that simultaneously predicts the occupancy and semantic labels of the environment. Compared with separate processing, SSC leverages the coupled nature of scene completion and semantic segmentation. Although this multitask integration can utilize complementarity and correlation between tasks, it also increases the training difficulty. To address this, in this letter, we propose a Semantic Decoupling based Semantic Scene Completion (SD-SSC) network from a single depth image. The semantic segmentation task is decoupled from the semantic scene completion task, and we use 2D and 3D semantic supervision to simplify the scene completion task and improve SSC performance. Specifically, our network first performs 2D semantic segmentation on the depth image and transforms features into 3D voxel space as semantic priors. Then, the 3D SSC is performed based on the voxel features and the flipped Truncated Signed Distance Field (f-TSDF). We use multi-scale 3D semantic supervision to further enhance the semantic information and fuse semantic and geometric features through the Planar Attention Fusion Module (PAFM) to obtain accurate SSC results. The proposed SD-SSC network achieves state-of-the-art performance on the NYU dataset (51.1% mIoU) and the NYUCAD dataset (61.9% mIoU) among all single depth-image based methods. It is even better than most RGB-D fusion-based SSC methods.
Keywords:
Semantic scene understanding
RGB-D perception
deep learning for visual perception
Journal
I
IF:
5.3
Papers:
1.7K
Citations:
3.9W

