Return
MapGlue: Multimodal remote sensing image matching
DOI:10.1016/j.isprsjprs.2026.07.007.png)
Abstract
En 中文
Multimodal remote sensing image (MRSI) matching is pivotal for cross-modal fusion, localization, and object detection, but it faces severe challenges due to geometric, radiometric, and viewpoint discrepancies across imaging modalities. Existing unimodal datasets lack scale and diversity, limiting deep learning solutions. This paper proposes MapGlue, a universal MRSI matching framework, and MapData, a large-scale multimodal dataset addressing these gaps. Our contributions are threefold. First, MapData, a globally diverse dataset spanning 233 sampling points, offers original images (7000 × 5,000 to 20,000 × 15,000 pixels). After rigorous cleaning, it provides 121,781 aligned electronic map–visible image pairs (512 × 512 pixels) with hybrid manual-automated ground truth, addressing the scarcity of scalable multimodal benchmarks. Second, the proposed MapGlue method constructs a saliency feature distribution map for keypoint extraction and integrates semantic information to accurately capture common invariant features. Furthermore, it utilizes a dual graph-guided mechanism to enable efficient global-to-local information interaction and robust feature enhancement for multimodal descriptors. Finally, extensive evaluations on MapData and five public datasets demonstrate MapGlue’s superiority in matching accuracy under complex conditions, outperforming state-of-the-art methods. Notably, MapGlue generalizes effectively to unseen modalities without retraining, highlighting its adaptability. This work addresses longstanding challenges in MRSI matching by combining scalable dataset construction with a robust, semantics-driven framework. The dataset and demo code are available at https://github.com/PeihaoWu/MapGlue .
Journal
IF:
12.2
Papers:
4.4K
Citations:
3.2W

