arrow
返回

Sampling for Big Data Profiling: A Survey

delete2020-01-01
delete14
delete
OA
AI
Z
Zhicheng Liu *
DOI:10.1109/ACCESS.2020.2988120delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Due to the development of internet technology and computer science, data is exploding at an exponential rate. Big data brings us new opportunities and challenges. On the one hand, we can analyze and mine big data to discover hidden information and get more potential value. On the other hand, the 5V characteristic of big data, especially Volume which means large amount of data, brings challenges to storage and processing. For some traditional data mining algorithms, machine learning algorithms and data profiling tasks, it is very difficult to handle such a large amount of data. The large amount of data is highly demanding hardware resources and time consuming. Sampling methods can effectively reduce the amount of data and help speed up data processing. Sampling technology has been widely used in big data context. Data profiling is the activity that finds metadata of data set and has many use cases, e.g., performing data profiling tasks on relational data, graph data, and time series data for anomaly detection and data repair. However, data profiling is computationally expensive, especially for large data sets. Hence this article focuses on researching sampling for data profiling tasks in big data context and investigates the application of sampling in different categories of data profiling. From the experimental results of these studies, the results got from the sampled data are close to or even exceed the results of the full amount of data. Therefore, sampling technology plays an important role in the era of big data, and we also have reason to believe that sampling technology will become an indispensable step in big data processing in the future.
Keyword:
Big Data
Data mining
Task analysis
Sampling methods
Metadata
Relational databases
Systematics
Big data
large amount
sampling
data profiling
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

T
tsinghua university
学者数:
11.9W
论文数: 10.0W
被引数: 137
U
University of Waterloo
学者数:
2.2W
论文数: 2.3W
被引数: 3.3W
引用论文

引用论文

err分享
err收藏
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Adaptive Control ofPhalaris arundinaceain Curtis Prairie
err2017-01-20
err0
PREAI
errMichael T. Healy; Isabel M. Rojas; Joy B. Zedler
err分享
err收藏
err分享
err收藏
Transition metal ion-doped In2O3nanocubes: investigation of their photocatalytic degradation activity under sunlight
err2021-01-01
err0
errOAAI
errVelayutham Shanmuganathan; Jayaraj Santhosh Kumar; Raman Pachaiappan; Paramasivam Thangadurai
err分享
err收藏
学者 查看更多内容