Return
Joint localization method via affine perception and multimodal fusion-based bird's-eye-view map registration
Y
W
C
D
Z
W
X
DOI:10.1117/1.JEI.35.2.023019.png)
Abstract
En 中文
In the field of autonomous driving, high-precision localization relies on effective registration between multimodal sensors and maps. However, due to issues such as misaligned multimodal features, viewpoint discrepancies, and depth estimation errors, existing bird's-eye-view (BEV) map (BEV-Map) joint localization methods often struggle to fully exploit the semantic priors embedded in maps and suffer from spatial misalignment across modalities. To address these challenges, we propose a BEV-Map joint localization method based on multimodal self-attention fusion and affine parameter optimization, aiming to achieve high-precision localization through enhanced registration accuracy. First, a multimodal self-attention fusion module is designed to strengthen the semantic complementarity between BEV and map features via bidirectional interaction. Second, an affine adaptive transformation module is introduced, which innovatively employs a parameter-separated prediction strategy to decouple the modeling of rotation, scaling, and translation, effectively mitigating cross-modal feature misalignment. Finally, a joint optimization scheme integrating geometric regularization and pose constraints is proposed to ensure that the predicted affine transformations maintain physical consistency while preserving accuracy. Experimental results demonstrate that the proposed method achieves a localization recall rate of 27.47% within 1 m and 94.53% within 10 m, outperforming existing state-of-the-art approaches. (c) 2026 SPIE and IS&T
Keywords:
multimodal
joint localization
registration
bird's-eye-view map
Journal
J
IF:
1
Papers:
109
Citations:
2.7K
