arrow
返回

CDFRS: A scalable sampling approach for efficient big data analysis

delete2024-07-01
delete1
delete
OA
AI
Y
Yongda Cai
D
Dingming Wu *
X
Xudong Sun
S
Siyue Wu
J
Joshua Zhexue Huang
DOI:10.1016/j.ipm.2024.103746delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
The sampling -based approximation method has demonstrated its potential in various domains such as machine learning, query processing, and data analysis. Most preceding sampling algorithms generate samples at the record level, making it impractical to apply them to very large datasets using a single machine. Even distributed solutions encounter efficiency issues when dealing with terabyte-scale datasets. In this paper, we introduce a scalable sampling approach named CDFRS, which can generate samples with a distribution -preserving guarantee from extensive datasets. CDFRS exhibits significantly improved speed compared to existing sampling algorithms when dealing with terabyte-scale datasets. We provide theoretical guarantees and empirical justifications, demonstrating that samples generated by the CDFRS approach maintain the distribution characteristics of the original dataset. Additionally, we propose a sample size determination algorithm, denoted as A 2 . Experiment results indicate that the running time of CDFRS shows at least an order of magnitude improvement over other distributed sampling methods. Notably, sampling a 10TB dataset using CDFRS only takes hundreds of seconds, while the compared method requires more than ten thousand seconds. In the context of big data analysis, including tasks such as classification and clustering, models trained with samples generated by CDFRS closely match those trained with the entire training set. Furthermore, the proposed A 2 algorithm efficiently determines an appropriate sample size compared with traditional methods.
Keyword:
Scalable sampling
Big data analysis
Random sample partition
Block-level sampling
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

I
Information Processing and Management
IF:
6.9
论文数:
5.2K
被引数:
1.4W

机构

S
shenzhen university
学者数:
4.6W
论文数: 3.4W
被引数: 72
引用论文

引用论文

Sampling for Big Data Profiling: A Survey
err2020-01-01
err14
errOAAI
errLiu, Zhicheng; Zhang, Aoqian
err分享
err收藏
Generative adversarial minority enlargement-A local linear over-sampling synthetic method
err2024-03-01
err3
PREAI
errWang, Ke; Zhou, Tongqing; Luo, Menghua; Li, Xionglve; Cai, Zhiping
err分享
err收藏
A scalable bootstrap for massive data
err2014-03-17
err286
errOAAI
errKleiner, Ariel; Talwalkar, Ameet; Sarkar, Purnamrita; Jordan, Michael I.
err分享
err收藏
Arsenic removal from aqueous solutions by adsorption using novel MIL-53(Fe) as a highly efficient adsorbent使用新型MIL-53(Fe) 作为高效吸附剂通过吸附从水溶液中去除砷
err2015-01-01
err0
PREAI
errTuan. A. Vu; Giang. H. Le; Canh. D. Dao; Lan. Q. Dang; Kien. T. Nguyen; Quang. K. Nguyen; Phuong. T. Dang; Hoa. T. K. Tran; Quang. T. Duong; Tuyen. V. Nguyen; Gun. D. Lee
err分享
err收藏
err分享
err收藏
学者 查看更多内容