arrow
Return

Applying Cluster Refinement to Improve Crowd-Based Data Duplicate Detection Approach

delete2019-01-01
delete0
delete
OA
AI
C
Charles Roland Haruna *
M
Mengshu Hou
R
Rui Xi
M
Moses Jojo Eghan
M
Michael Y. Kpiebaareh
L
Lawrence Tandoh
B
Barbie Eghan-Yartel
M
Maame G. Asante-Mensah
DOI:10.1109/ACCESS.2019.2920667delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
In this paper, we present an extension on a hybrid-based deduplication technique in entity reconciliation (ER), by proposing an algorithm that builds clusters upon receiving a pre-specified K number of clusters, and second developing a crowd-based procedure for refining the results of the clusters produced after the clustering generation phases. With the clusters refined, we aim to minimize the cost metric Lambda'(R) of the solitary and compound cluster generation algorithms, to achieve an improved and efficient deduplication method, to have an increase in accuracy in identifying duplicate records, and finally, further reduce the crowdsourcing overheads incurred. In this paper, in the experiments, we made use of three datasets commonly known to hybrid-based deduplication such as paper, product, and restaurant. The performance results and evaluations demonstrate clear superiority to the methods compared with our work offering low-crowdsourcing cost and high accuracy of deduplication, as well as better deduplication efficiency due to the clusters being refined.
Keywords:
Cluster refinement
minimization approach
triangular split and merger operations
entity reconciliation
crowdsourcing
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

U
university of cape coast
Scholars:
2.8K
Papers: 1.6K
Citations: 3