arrow
返回

High-Value Token-Blocking: Efficient Blocking Method for Record Linkage

delete2021-07-21
delete2
delete
OA
AI
K
Kevin O’Hare *
A
Anna Jurek-Loughrey
C
Cassio P. de Campos
DOI:10.1145/3450527delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Data integration is an important component of Big Data analytics. One of the key challenges in data integration is record linkage, that is, matching records that represent the same real-world entity. Because of computational costs, methods referred to as blocking are employed as a part of the record linkage pipeline in order to reduce the number of comparisons among records. In the past decade, a range of blocking techniques have been proposed. Real-world applications require approaches that can handle heterogeneous data sources and do not rely on labelled data. We propose high-value token-blocking (HVTB), a simple and efficient approach for blocking that is unsupervised and schema-agnostic, based on a crafted use of Term Frequency-Inverse Document Frequency. We compare HVTB with multiple methods and over a range of datasets, including a novel unstructured dataset composed of titles and abstracts of scientific papers. We thoroughly discuss results in terms of accuracy, use of computational resources, and different characteristics of datasets and records. The simplicity of HVTB yields fast computations and does not harm its accuracy when compared with existing approaches. It is shown to be significantly superior to other methods, suggesting that simpler methods for blocking should be considered before resorting to more sophisticated methods.
Keyword:
Blocking
record linkage
entity resolution
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

ACM Transactions on Knowledge Discovery from Data 封面图
ACM Transactions on Knowledge Discovery from Data
IF:
4.8
论文数:
1.3K
被引数:
4.4K

机构

Q
Queen's University Belfast
学者数:
1.6W
论文数: 1.7W
被引数: 2.5W
E
Eindhoven University of Technology
学者数:
1.6W
论文数: 1.5W
被引数: 2.2W