arrow
Return

Scalable Iterative Graph Duplicate Detection

delete2012-11-01
delete16
delete
OA
AI
M
Melanie Herschel *
F
Felix Naumann
DOI:10.1109/TKDE.2011.99delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Duplicate detection determines different representations of real-world objects in a database. Recent research has considered the use of relationships among object representations to improve duplicate detection. In the general case where relationships form a graph, research has mainly focused on duplicate detection quality/effectiveness. Scalability has been neglected so far, even though it is crucial for large real-world duplicate detection tasks. We scale-up duplicate detection in graph data (DDG) to large amounts of data and pairwise comparisons, using the support of a relational database management system. To this end, we first present a framework that generalizes the DDG process. We then present algorithms to scale DDG in space (amount of data processed with bounded main memory) and in time. Finally, we extend our framework to allow batched and parallel DDG, thus further improving efficiency. Experiments on data of up to two orders of magnitude larger than data considered so far in DDG show that our methods achieve the goal of scaling DDG to large volumes of data.
Keywords:
Duplicate detection
data cleaning
data integration
record linkage
entity resolution
scalability
parallelization
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Knowledge and Data Engineering cover
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
Papers:
6.7K
Citations:
3.2W

Organization

E
eberhard karls university of tubingen
Scholars:
3.3W
Papers: 2.5W
Citations: 38
Zuse Institute Berlin cover
Zuse Institute Berlin
Scholars:
420
Papers: 350
Citations: 367