Return
Enhancing Point Cloud Feature Utilization for 3D Object Detection
DOI:10.1109/LSP.2025.3626268.png)
Abstract
En 中文
Traditional voxelization methods use many artificial components to handle sparse and uneven point clouds, which may lead to spatial discretization distortion. In addition, the fixed convolution size limits the detector's ability to capture feature correlations, resulting in low feature utilization. To adderss this challenge, this paper proposes a two-stage LiDAR 3D object detector, Pillar-CT3D++. First, it introduces a self-attention encoding module that uses a Transformer to process voxelized point cloud information. By capturing the position information of point cloud objects within voxels through multi-dimensional position encoding and an encoder-decoder mechanism, the quality of feature suggestions is improved. Second, a feature weight-aware module is proposed, which designs multi-scale pooling and group attention mechanisms in the 2D backbone network to adaptively recalibrate the feature responses of channels, generating more representative feature information. The detection performance of Pillar-CT3D++ is validated through experiments on the KITTI dataset and compared with existing detectors. Specifically, on the KITTI dataset, the proposed model outperforms the baseline CT3D by 0.43%, 1.04%, and 0.87% in detecting simple, medium, and difficult-level objects, respectively.
Keywords:
Feature extraction
Point cloud compression
Encoding
Three-dimensional displays
Object detection
Convolution
Vectors
Detectors
Decoding
Attention mechanisms
3D object detection
autonomous driving
LiDAR
point cloud
Journal
I
IF:
3.9
Papers:
610
Citations:
0

