arrow
Return

Efficient and Effective Duplicate Detection in Hierarchical Data

delete2013-05-01
delete16
delete
OA
AI
L
Luís Leitão *
P
Pável Calado
M
Melanie Herschel
DOI:10.1109/TKDE.2012.60delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Although there is a long line of work on identifying duplicates in relational data, only a few solutions focus on duplicate detection in more complex hierarchical structures, like XML data. In this paper, we present a novel method for XML duplicate detection, called XMLDup. XMLDup uses a Bayesian network to determine the probability of two XML elements being duplicates, considering not only the information within the elements, but also the way that information is structured. In addition, to improve the efficiency of the network evaluation, a novel pruning strategy, capable of significant gains over the unoptimized version of the algorithm, is presented. Through experiments, we show that our algorithm is able to achieve high precision and recall scores in several data sets. XMLDup is also able to outperform another state-of-the-art duplicate detection solution, both in terms of efficiency and of effectiveness.
Keywords:
Duplicate detection
record linkage
entity resolution
XML
Bayesian networks
data cleaning
optimization
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Knowledge and Data Engineering cover
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
Papers:
6.7K
Citations:
3.2W

Organization

U
universidade de lisboa
Scholars:
3.4W
Papers: 3.1W
Citations: 29
U
Universite Paris Saclay
Scholars:
7.3W
Papers: 5.3W
Citations: 540