返回
Duplicate record detection: A survey
DOI:10.1109/TKDE.2007.250581.png)
摘要
En 中文
Often, in the real world, entities have two or more representations in databases. Duplicate records do not share a common key and/or they contain errors that make duplicate matching a difficult task. Errors are introduced as the result of transcription errors, incomplete information, lack of standard formats, or any combination of these factors. In this paper, we present a thorough analysis of the literature on duplicate record detection. We cover similarity metrics that are commonly used to detect similar field entries, and we present an extensive set of duplicate detection algorithms that can detect approximately duplicate records in a database. We also cover multiple techniques for improving the efficiency and scalability of approximate duplicate detection algorithms. We conclude with coverage of existing tools and with a brief discussion of the big open problems in the area.
Keyword:
duplicate detection
data cleaning
data integration
record linkage
data deduplication
instance identification
database hardening
name matching
identity uncertainty
entity resolution
fuzzy duplicate detection
entity matching
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
10.4
论文数:
6.8K
被引数:
3.2W
机构
暂无机构信息
引用论文
Characterization of polymer matrix and low melting point solder for anisotropic conductive film各向异性导电膜用聚合物基体及低熔点焊料的表征

