Return
Probabilistic Record Linkage Using Pretrained Text Embeddings
DOI:10.1017/pan.2025.10016.png)
Abstract
En 中文
Pretrained text embeddings are a fast and scalable method for determining whether two texts have similar meaning; capturing not only lexical similarity; but semantic similarity as well. In this article; I show how to incorporate these measures into a probabilistic record linkage procedure that yields considerable improvements in both precision and recall over existing methods. The procedure even allows researchers to link datasets across different languages. I validate the approach with a series of political science applications; and provide open-source statistical software for researchers to efficiently implement the proposed method.
Keywords:
probabilistic record linkage
fuzzy string matching
embeddings
large language models (LLMs)
GPT-3
GPT-4
active learning
text-as-data
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

