arrow
返回

A distributed data management system to support large-scale data analysis

delete2019-02-01
delete16
PRE
AI
T
Tamer Z. Emara
J
Joshua Zhexue Huang *
DOI:10.1016/j.jss.2018.11.007delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Distributed data management is a key technology to enable efficient massive data processing and analysis in cluster-computing environments. Specifically, in environments where the data volumes are beyond the system capabilities, big data files are required to be summarized by representative samples with the same statistical properties as the whole dataset. This paper proposes a big data management system (BDMS) based on distributed random sample data blocks. It presents a high-level architecture design of the BDMS which extends the current distributed file systems. This system offers certain functionalities for block-level management such as statistically-aware data partitioning, data blocks organization, and data blocks selection. This paper also presents a round-random partitioning scheme to represent a big dataset as a set of non-overlapping data blocks; each block is a random sample of the whole dataset. Based on the presented scheme, two algorithms are introduced as an implementation strategy to convert the HDFS blocks of a big file into a set of random sample data blocks which is also stored in HDFS. The experimental results show that the execution time of partitioning operation is acceptable in the real applications because this operation is only performed once on each input data file. (C) 2018 Elsevier Inc. All rights reserved.
Keyword:
Big data
Distributed and parallel processing
Random sample partition
Randomness
Data management
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Journal of Systems and Software 封面图
Journal of Systems and Software
IF:
4.1
论文数:
5.5K
被引数:
8.4K

机构

S
shenzhen university
学者数:
4.6W
论文数: 3.4W
被引数: 72
引用论文

引用论文

MapReduce and Parallel DBMSs: Friends or Foes?MapReduce和并行dbms: 朋友还是敌人?
err2010-01-01
err226
PREAI
errStonebraker, Michael; Abadi, Daniel; Dewitt, David J.; Madden, Sam; Paulson, Erik; Pavlo, Andrew; Rasin, Alexander
err分享
err收藏
A scalable bootstrap for massive data
err2014-03-17
err286
errOAAI
errKleiner, Ariel; Talwalkar, Ameet; Sarkar, Purnamrita; Jordan, Michael I.
err分享
err收藏
Apache Spark: A Unified Engine for Big Data ProcessingApache Spark: 用于大数据处理的统一引擎
err2016-10-28
err1.7K
PREAI
errZaharia, Matei; Xin, Reynold S.; Wendell, Patrick; Das, Tathagata; Armbrust, Michael; Dave, Ankur; Meng, Xiangrui; Rosen, Josh; Venkataraman, Shivaram; Franklin, Michael J.; Ghodsi, Ali; Gonzalez, Joseph; Shenker, Scott; Stoica, Ion
err分享
err收藏
没有更多内容