Return
Fadet: a fusion-aware 3D detection network with cascaded feature enhancement for small object detection in autonomous driving
C
S
W
DOI:10.1007/s00530-026-02585-3.png)
Abstract
En 中文
Small object detection is one of the key challenges in autonomous driving perception, where both accuracy and real-time efficiency are critical to ensuring driving safety. Although multimodal fusion has become the mainstream paradigm, existing approaches often neglect the inherent sparsity of LiDAR point clouds and the alignment deviations between point clouds and images. As a result, the already weak features of small objects are further diluted during fusion, which constrains performance gains. To address these limitations, we propose a cascade optimization framework that systematically enhances feature representation and refines multimodal fusion. Specifically, a Channel-Pillar Hybrid Attention Module (CPHAM) is introduced to alleviate the insufficient feature representation caused by point cloud sparsity. By sequentially recalibrating feature importance along the channel and pillar dimensions, CPHAM adaptively highlights discriminative regions and provides more accurate features for subsequent fusion. Furthermore, we introduce the Multi-Scale Contextual Fusion Module (MSCF) to reduce alignment bias. MSCF strengthens geometric connectivity, aligns multimodal features more precisely, and integrates PB-FPN for multi-scale refinement to preserve and enhance small-object features during fusion. Together, the two modules work progressively to prevent small-object features from being weakened. Extensive experiments on the nuScenes dataset demonstrate that our method achieves 75.2% mAP and 76.8% NDS, validating its effectiveness and superiority.
Keywords:
Multi-modal small object detection
Attention mechanism
Feature enhancement
Contextual fusion
Journal
IF:
3.1
Papers:
2.7K
Citations:
2.7K
