arrow
Return

SuperMapNet for long-range and high-accuracy vectorized HD map construction

delete2026-01-20
delete0
PRE
AI
R
Ruqin Zhou
C
Chenguang Dai
W
Wanshou Jiang
Y
Yongsheng Zhang
Z
Zhenchao Zhang
姜三 (San Jiang) *
DOI:10.1016/j.isprsjprs.2026.01.023delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Vectorized high-definition (HD) map construction is formulated as the task of classifying and localizing typical map elements based on features in a bird’s-eye view (BEV). This is essential for autonomous driving systems, providing interpretable environmental structured representations for decision and planning. Remarkable work has been achieved in recent years, but several major issues remain: (1) in the generation of the BEV features, single modality methods suffer from limited perception capability and range, while existing multi-modal fusion approaches underutilize cross-modal synergies and fail to resolve spatial disparities between modalities, resulting in misaligned BEV features with holes; (2) in the classification and localization of map elements, existing methods heavily rely on point-level modeling information while neglecting the information between elements and between point and element, leading to low accuracy with erroneous shapes and element entanglement. To address these limitations, we propose SuperMapNet, a multi-modal framework designed for long-range and high-accuracy vectorized HD map construction. This framework uses both camera images and LiDAR point clouds as input. It first tightly couples semantic information from camera images and geometric information from LiDAR point clouds by a cross-attention based synergy enhancement module and a flow-based disparity alignment module for long-range BEV feature generation. Subsequently, local information acquired by point queries and global information acquired by element queries are tightly coupled by three-level interactions for high-accuracy classification and localization, where Point2Point interaction captures local geometric consistency between points of the same element, Element2Element interaction learns global semantic relationships between elements, and Point2Element interaction complement element information for its constituent points. Experiments on the nuScenes and Argoverse2 datasets demonstrate high accuracy, surpassing previous state-of-the-art methods (SOTAs) by 14.9%/8.8% and 18.5%/3.1% mAP under the hard/easy settings, respectively, even over the double perception ranges (up to 120 m in the X-axis and 60 m in the Y-axis). The code is made publicly available at https://github.com/zhouruqin/SuperMapNet .

Journal

ISPRS Journal of Photogrammetry and Remote Sensing cover
ISPRS Journal of Photogrammetry and Remote Sensing
IF:
12.2
Papers:
4.4K
Citations:
3.2W

Organization

S
Shenzhen University
Scholars:
4.0K
Papers: 1.7K
Citations: 5.4W
I
Information Engineering University
Scholars:
484
Papers: 161
Citations: 0
W
Wuhan University
Scholars:
5.0K
Papers: 1.7K
Citations: 10.0W
researcher View more organizations