arrow
Return

Quality-Based Online Data Reconciliation

delete2016-02-01
delete2
PRE
AI
A
Asma Abboura *
S
Soror Sahri
L
Latifa Baba-Hamed *
M
Mourad Ouziri
S
Salima Benbernou *
DOI:10.1145/2806888delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
One of the main challenges in data matching and data cleaning, in highly integrated systems, is duplicates detection. While the literature abounds of approaches detecting duplicates corresponding to the same real world entity, most of these approaches tend to eliminate duplicates (wrong information) from the sources, hence leading to what is called data repair. In this article, we propose a framework that automatically detects duplicates at query time and effectively identifies the consistent version of the data, while keeping inconsistent data in the sources. Our framework uses matching dependencies (MDs) to detect duplicates through the concept of data reconciliation rules (DRR) and conditional function dependencies (CFDs) to assess the quality of different attribute values. We also build a duplicate reconciliation index (DRI), based on clusters of duplicates detected by a set of DRRs to speed up the online data reconciliation process. Our experiments of a real-world data collection show the efficiency and effectiveness of our framework.
Keywords:
Duplicates
data reconciliation
data quality rules
source quality
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

ACM Transactions on Internet Technology cover
ACM Transactions on Internet Technology
IF:
4.1
Papers:
896
Citations:
1.9K

Organization

U
Universite Paris Cite
Scholars:
8.9W
Papers: 6.3W
Citations: 604