arrow
返回

SuperMapNet for long-range and high-accuracy vectorized HD map construction

delete2026-01-20
delete0
PRE
AI
R
Ruqin Zhou
C
Chenguang Dai
W
Wanshou Jiang
Y
Yongsheng Zhang
Z
Zhenchao Zhang
姜
姜三 (San Jiang) *
DOI:10.1016/j.isprsjprs.2026.01.023delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Vectorized high-definition (HD) map construction is formulated as the task of classifying and localizing typical map elements based on features in a bird’s-eye view (BEV). This is essential for autonomous driving systems, providing interpretable environmental structured representations for decision and planning. Remarkable work has been achieved in recent years, but several major issues remain: (1) in the generation of the BEV features, single modality methods suffer from limited perception capability and range, while existing multi-modal fusion approaches underutilize cross-modal synergies and fail to resolve spatial disparities between modalities, resulting in misaligned BEV features with holes; (2) in the classification and localization of map elements, existing methods heavily rely on point-level modeling information while neglecting the information between elements and between point and element, leading to low accuracy with erroneous shapes and element entanglement. To address these limitations, we propose SuperMapNet, a multi-modal framework designed for long-range and high-accuracy vectorized HD map construction. This framework uses both camera images and LiDAR point clouds as input. It first tightly couples semantic information from camera images and geometric information from LiDAR point clouds by a cross-attention based synergy enhancement module and a flow-based disparity alignment module for long-range BEV feature generation. Subsequently, local information acquired by point queries and global information acquired by element queries are tightly coupled by three-level interactions for high-accuracy classification and localization, where Point2Point interaction captures local geometric consistency between points of the same element, Element2Element interaction learns global semantic relationships between elements, and Point2Element interaction complement element information for its constituent points. Experiments on the nuScenes and Argoverse2 datasets demonstrate high accuracy, surpassing previous state-of-the-art methods (SOTAs) by 14.9%/8.8% and 18.5%/3.1% mAP under the hard/easy settings, respectively, even over the double perception ranges (up to 120 m in the X-axis and 60 m in the Y-axis). The code is made publicly available at https://github.com/zhouruqin/SuperMapNet .

期刊

ISPRS Journal of Photogrammetry and Remote Sensing 封面图
ISPRS Journal of Photogrammetry and Remote Sensing
IF:
12.2
论文数:
4.4K
被引数:
3.2W

机构

S
Shenzhen University
学者数:
4.0K
论文数: 1.7K
被引数: 5.4W
I
Information Engineering University
学者数:
484
论文数: 161
被引数: 0
W
Wuhan University
学者数:
5.0K
论文数: 1.7K
被引数: 10.0W
学者 查看更多机构
引用论文

引用论文

暂无论文信息