Return
Decoding nested entities from classical Chinese with LLMs
DOI:10.1080/13658816.2025.2583431.png)
Abstract
En 中文
Historical texts, particularly classical Chinese texts, contain a variety of domain-specific entities with latent structures that remain to be exploited. This study extracted and analyzed nested entities from large-scale historical texts, using Chinese local gazetteers as a representative dataset. We designed a data sampling strategy integrated with the multi-objective optimization algorithm NSGA-II to address entity imbalances in low-resource and domain-specific areas. Then we utilized two large language models, Qwen 2.5-7B and Xunzi-Qwen2-7B, with LoRA tuning and prompt learning for entity extraction. We compared this method with three non-generative approaches - GlobalPointer, Biaffine, and TPLinker - and used multiple metrics, including micro-average, macro-average, exact match, position match, and depth match, to evaluate the model's performance. Next, we assessed the cross-dataset generalization of the model on two additional historical datasets and confirmed its robustness across domains. The experimental results illustrated the feasibility and effectiveness of data sampling and Large Language Model-based approaches for processing Chinese historical texts. Furthermore, we investigated the intra-entity patterns embedded in nested entities using the Apriori algorithm. This on-going research enriches existing geographical repositories by providing valuable toponym information, and offers insightful reference for other areas of historical text processing.
Keywords:
Local gazetteers
nested entity
LLM
Journal
IF:
5.1
Papers:
2.7K
Citations:
9.3K

