arrow
返回

Generating Name-Like Vectors for Testing Large-Scale Entity Resolution

delete2021-01-01
delete0
delete
OA
AI
S
Samudra Herath *
M
Matthew Roughan
G
Gary Glonek
DOI:10.1109/ACCESS.2021.3122451delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Entity resolution (ER), the problem of identifying and linking records that belong to the same real-world entities in structured and unstructured data, is a primary task in data integration. Accurate and efficient ER has a major practical impact on various applications across commercial, security and scientific domains. Recently, scalable ER techniques have received enormous attention with the increasing need to combine large-scale datasets. The shortage of training and ground truth data impedes the development and testing of ER algorithms. Good public datasets, especially those containing personal information, are restricted in this area and usually small in size. Due to privacy and confidential issues, testing algorithms or techniques with real datasets is challenging in ER research. Simulation is one technique for generating synthetic datasets that have characteristics similar to those of real data for testing algorithms. Many existing simulation tools in ER lack support for generating large-scale data and have problems in complexity, scalability, and limitations of resampling. In our work, we propose a simple, inexpensive, and fast synthetic data generation tool. Our tool only generates entity names in the first stage, but these are commonly used as identification keys in ER algorithms. We avoid the detail-level simulation of entity names using a simple vector representation that delivers simplicity and efficiency. In this paper, we discuss how to simulate simple vectors that approximate the properties of entity names. We describe the overall construction of the tool based on data analysis of a namespace that contains entity names collected from the actual environment.
Keyword:
Data models
Erbium
Databases
Tools
Big Data
Testing
Numerical models
Entity resolution
data integration
data linkage
data matching
information systems
large-scale synthetic data
record linkage

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

U
University of Adelaide
学者数:
2.3W
论文数: 2.4W
被引数: 4.2W
引用论文

引用论文

err分享
err收藏
err2004-01-01
err0
errOAAI
errDennis C Turk; Robert H Dworkin
err分享
err收藏
Multidimensional scaling
err2012-10-08
err329
errOAAI
errHout, Michael C.; Papesh, Megan H.; Goldinger, Stephen D.
err分享
err收藏
Multidimensional Scaling by Majorization: A Review
err2016-01-01
err22
errOAAI
errGroenen, Patrick J. F.; van de Velden, Michel
err分享
err收藏
Electrical resistivity and7Li Knight shift of liquid Li-Si alloys
err1999-01-01
err0
PREAI
errJ A Meijer; C van der Marel; P Kuiper; W van der Lugt
err分享
err收藏
err分享
err收藏
An Overview of End-to-End Entity Resolution for Big Data
err2020-12-06
err84
errOAAI
errChristophides, Vassilis; Efthymiou, Vasilis; Palpanas, Themis; Papadakis, George; Stefanidis, Kostas
err分享
err收藏
学者 查看更多内容