arrow
Return

Probabilistic Record Linkage Using Pretrained Text Embeddings

delete2026-03-28
delete0
delete
OA
AI
J
Joseph T. Ornstein
DOI:10.1017/pan.2025.10016delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Pretrained text embeddings are a fast and scalable method for determining whether two texts have similar meaning; capturing not only lexical similarity; but semantic similarity as well. In this article; I show how to incorporate these measures into a probabilistic record linkage procedure that yields considerable improvements in both precision and recall over existing methods. The procedure even allows researchers to link datasets across different languages. I validate the approach with a series of political science applications; and provide open-source statistical software for researchers to efficiently implement the proposed method.
Keywords:
probabilistic record linkage
fuzzy string matching
embeddings
large language models (LLMs)
GPT-3
GPT-4
active learning
text-as-data
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Political Analysis cover
Political Analysis
IF:
5.4
Papers:
74
Citations:
6.7K

Organization

U
University of Georgia
Scholars:
1.5W
Papers: 1.2W
Citations: 2.9W