1
Return

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization

delete2026-07-27
delete0
PRE
AI
W
Wenping Yin
Z
Ziqi Liu
N
Naixia Mou
W
Weijia Li
D
Danfeng Hong
H
Hao Li
DOI:10.1109/tgrs.2026.3717084delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Social media imagery (SMI) provides timely and fine-grained ground perspectives that are valuable for situational awareness and emergency response. Unlike satellite or aerial imagery, SMI can capture disaster impacts and ground-level conditions in a timely manner. However, geographic references in SMI are often vague or ambiguous, making accurate geolocalization challenging. To address this issue, we propose Disaster TD, a disaster toponym disambiguation framework that integrates multimodal large language models (MLLMs)-based semantic reasoning with cross-view geolocalization. First, MLLMs extract toponyms and generate candidate geolocations from noisy textual inputs. Then, cross-view matching between SMI, remote sensing imagery (RSI), and optionally street-view imagery (SVI) is used to verify and refine these candidate results. A Vision Transformer (ViT)-based visual foundation model, DINOv2, is used to bridge the domain gap between overhead and ground-level imagery. We evaluate DisasterTD on the Hurricane Harvey dataset, where SMI is augmented with collected RSI and SVI to construct a cross-view benchmark for disaster geolocalization. The dataset is divided into four categories based on toponym clarity and ambiguity, allowing a fine-grained performance analysis across scenarios. Results show that DisasterTD consistently outperforms MLLM-only and cross-view-only baselines without disambiguation, achieving geolocalization accuracies of 71.62% within 1000 m, 62.36% within 500 m, 57.99% within 250 m, 52.09% within 100 m, and 47.01% within 50 m, while reducing the mean and median errors to 11.33 and 0.68 km, respectively. The largest improvements appear in ambiguous toponyms, where semantic reasoning with cross-view evidence reduces candidate dispersion and errors. These findings demonstrate the effectiveness of integrating MLLM-based candidate generation with cross-view verification for fine-grained disaster geolocalization.
Keywords:
Cross-view
disaster response
geolocalization
multimodal LLM
toponym disambiguation

Journal

IEEE Transactions on Geoscience and Remote Sensing cover
IEEE Transactions on Geoscience and Remote Sensing
IF:
8.6
Papers:
2.1W
Citations:
10.7W

Organization

S
shandong university of science and technology
Scholars:
1.8K
Papers: 578
Citations: 0
T
tsinghua university
Scholars:
11.5W
Papers: 9.9W
Citations: 137
S
Southeast University
Scholars:
1.8W
Papers: 7.6K
Citations: 480
N
National University of Singapore
Scholars:
7.4W
Papers: 6.4W
Citations: 11.4W
W
wuhan university
Scholars:
7.8W
Papers: 5.7W
Citations: 70
Cited Papers

Cited Papers

Citing Papers

Citing Papers