Return
A geographic entity recognition method utilizing temporal active learning and large language models
DOI:10.1080/10095020.2026.2657656.png)
Abstract
En 中文
Employing active learning is a viable approach to reducing the human effort required to complete geographic entity recognition tasks using deep learning methods. Nevertheless, current sampling strategies often only capture one sample characteristic (e.g. information content) and disregard the temporal dynamics of the sample feature across different iterations. Moreover, active learning methods alone can only alleviate workload at the sample selection level. To address these issues, we propose a novel active learning framework that integrates temporal features into sampling and leverages a large language model (LLM) to enhance annotation efficiency and the precision of model prediction. Specifically, we propose a mixed sampling strategy based on temporal and semantic similarity, comprising three innovative indicators: temporal instability, dynamic variance entropy, and semantic similarity. We propose an approach to utilizing the LLM Qwen2.5-7B-Instruct for machine annotation and an expert system for annotation correction, further reducing the workload of geographic entity recognition tasks. Experiments demonstrate that our proposed sampling strategies outperform random sampling, with “temporal instability” exhibiting optimal performance and robustness. Furthermore, combining machine annotation with expert systems can save approximately 27% of annotation time. Our work provides a practical solution for designing sampling strategies with temporal features and fully leveraging LLM for automatic annotation to reduce the annotation cost of geographic entity recognition tasks.
Keywords:
Active learning
temporal characteristics
large language models (LLMs)
geographic entity recognition
BiLSTM-CRF
Journal
G
IF:
5.5
Papers:
836
Citations:
2.4K

