Return
Tokenization and representation learning of contour data using bidirectional encoder representations from transformers (BERT)
DOI:10.1080/13658816.2026.2643000.png)
Abstract
En 中文
Large language models have demonstrated a transformative capability in advancing intelligence across multiple domains. In geoscience, vector map data serves as a fundamental data source but faces challenges in information mining. Contour data constitute a graphic language for terrain representation and exhibit a sequential structure analogous to text in natural language processing, making integration with transformer-based models (TBMs) possible. This study adopts a TBM to enhance the representation and information mining of contour data. Similar to a text corpus composed of words and sentences, we tokenized contour data from 40 regions into a Contour Corpus composed of tokens and sequences. Then, a Bidirectional Encoder Representations from Transformers (BERT) model was pre-trained on this corpus, termed Contour-BERT. Intrinsic evaluation metrics and representation analysis confirmed that the model could capture meaningful representations. Finally, the model was applied to three downstream tasks: terrain feature recognition, contour pattern classification, and geomorphological unit classification. The experimental results demonstrated that Contour-BERT outperformed the other methods. This study proposes a novel framework for the tokenization and representation learning of contour data, offering an alternative avenue for integrating vector map data with AI technologies.
Keywords:
Contour data
tokenization
representation learning
bidirectional encoder representations from transformers (BERT)
information mining
Journal
IF:
5.1
Papers:
2.7K
Citations:
9.3K

