Return
An improved multi-sources toponymic matching method by integrated encoding of textual and spatial features
M
T
X
M
S
DOI:10.1080/15230406.2026.2624404.png)
Abstract
En 中文
Multi-source toponymic matching is the process of consistently discriminating between geometric and attribute information of toponyms, so as to obtain more accurate, complete, and high-quality data. Attribute information of toponyms is expressed by natural language symbols, while geometric information of toponyms is expressed by coordinate values. This difference in expression system and expression mode leads to challenges in interpreting the semantic and spatial features of toponyms. At the present stage, geographic entity matching mainly adopts similarity index calculation method to compare toponymic features. However, these indices rely on subjective judgment thresholds, making it difficult to fully capture and analyze the semantic and spatial features of toponyms. Therefore, we propose a method for deep feature extraction of toponyms, embedding textual and spatial features into a unified high-dimensional vector representation space and using an attention mechanism to capture the intersection features between attributes to obtain the vector representation. Evaluated on GeoNames and OSM, our method achieves 98.20% accuracy and 0.9823 F1-score, significantly outperforming models based on character distances similarity and machine learning classifiers. Results demonstrate effective that this method achieves an effective integration of semantic and spatial features, enabling robust toponymic matching.
Keywords:
Text embedding
location encoding
toponymic matching
representing learning
Journal
IF:
2.4
Papers:
103
Citations:
1.5K
