Return
LSON-IP: Lightweight Sparse Occupancy Network for Instance Perception
DOI:10.3390/wevj17010031.png)
Abstract
En 中文
The high computational demand of dense voxel representations severely limits current vision-centric 3D semantic occupancy prediction methods, despite their capacity for granular scene understanding. This challenge is particularly acute in safety-critical applications like autonomous driving, where accurately perceiving dynamic instances often takes precedence over capturing the static background. This paper challenges the paradigm of dense prediction for such instance-focused tasks. We introduce the LSON-IP, a framework that strategically avoids the computational expense of dense 3D grids. LSON-IP operates on a sparse set of 3D instance queries, which are initialized directly from multi-view 2D images. These queries are then refined by our novel Sparse Instance Aggregator (SIA), an attention-based module. The SIA incorporates rich multi-view features while simultaneously modeling inter-query relationships to construct coherent object representations. Furthermore, to obviate the need for costly 3D annotations, we pioneer a Differentiable Sparse Rendering (DSR) technique. DSR innovatively defines a continuous field from the sparse voxel output, establishing a differentiable bridge between our sparse 3D representation and 2D supervision signals through volume rendering. Extensive experiments on major autonomous driving benchmarks, including SemanticKITTI and nuScenes, validate our approach. LSON-IP achieves strong performance on key dynamic instance categories and competitive overall semantic completion, all while reducing computational overhead by over 60% compared to dense baselines. Our work thus paves the way for efficient, high-fidelity instance-aware 3D perception.
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

