返回
SCRTN: Enhancing multi-modal 3D object detection in complex environments
DOI:10.1016/j.patcog.2026.113068.png)
摘要
En 中文
• A multi-modal 3D object detection framework, the SCRTN, is proposed to fuse 3D point cloud data with two-dimensional (2D) image data, which enhances recognition performance of 3D objects in complex environments. • The ResTransfusion feature fusion technique is used to strengthen the global information connection between the point cloud features and enhanced point clouds, which improves the synergy between semantic and shape features, thus improving the model’s depth optimization capability. • A far-reaching voxel preservation sampling strategy is designed and combined with sparse convolutional technology to increase the specificity of feature extraction in 3D data, which increases the efficiency of 3D object detection. • The results of the extensive experiments demonstrate that the proposed method can achieve the state-of-the-art performance on the Hard KITTI dataset, with an accuracy of 85.75 and a mean average precision (mAP) of 89.67. Moreover, the proposed model also has excellent performance on the Nuscenes dataset and Waymo Open Dataset.
期刊
IF:
7.6
论文数:
1.3W
被引数:
4.5W
机构
引用论文
Beyond masking: Demystifying token-based pre-training for vision transformers超越掩码:解析基于token的视觉Transformer预训练
PATTERN RECOGNITION
IF7.6
LiDARCapV2: 3D human pose estimation with human-object interaction from LiDAR point cloudsLiDARCapV2: 基于激光雷达点云的人-物交互的3D人体姿态估计
PATTERN RECOGNITION
IF7.6
没有更多内容

