arrow
返回

Hadoop Data Reduction Framework: Applying Data Reduction at the DFS Layer

delete2021-01-01
delete1
delete
OA
AI
R
Ryan Nathanael Soenjoto Widodo *
H
Hirotake Abe
K
Kazuhiko Kato
DOI:10.1109/ACCESS.2021.3127499delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Big-data processing systems such as Hadoop, which usually utilize distributed file systems (DFSs), require data reduction schemes to maximize storage space efficiency. These schemes have different tradeoffs, and there are no all-purpose schemes applicable to all data. Users must select a suitable scheme in accordance with their data. To accommodate this requirement, application software or file system (FS) have a fixed selection of these schemes. However, these provided schemes are insufficient for all data types, and when novel schemes emerge, extending the selection can be problematic. If the source code of the application or FS is available, the source code could potentially be extended with extensive labor, but could be virtually impossible without the code maintainers' assistance. If the source code is unavailable, there is no way to tackle the problem. This paper proposes an unexplored solution through a modular DFS design that eases data reduction scheme usage through existing programming techniques. The advantages of this presented approach are threefold. First, adding new schemes is easy and they are transparent to the application code requiring no extensions to it. Second, the modular structure requires minimal modification to the existing DFSs and performance overhead. Third, users can compile schemes separately from the DFS without the FS or DFS source code. To demonstrate the design's effectiveness, we implemented it by minimally extending the Hadoop DFS (HDFS) and named it the Hadoop Data Reduction Framework (HDRF). We designed HDRF to work with minimal overhead and tested it extensively. Experimental results indicate that it has negligible overhead over existing approaches. In a number of cases, it can offer up to 48.96% higher throughput while achieving the best result in storage reduction within our tested setups because of the incorporated data reduction schemes.
Keyword:
Codes
Software
Dictionaries
Redundancy
Libraries
Big Data
File systems
Data compression
data deduplication
distributed file system
Hadoop
HDFS

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

U
University of Tsukuba
学者数:
1.8W
论文数: 1.5W
被引数: 1.7W
引用论文

引用论文

Impact of CYP3A5 Genetic Polymorphism on Intrapatient Variability of Tacrolimus Exposure in Chinese Kidney Transplant Recipients
err2019-07-01
err0
errOAAI
errChi Yuen Cheung; Koon Ming Chan; Yuen Ting Wong; Wai Leung Chak; Otto Bekers; Johannes P. van Hooff
err分享
err收藏
A Survey of Secure Data Deduplication Schemes for Cloud Storage Systems
err2017-01-09
err87
PREAI
errShin, Youngjoo; Koo, Dongyoung; Hur, Junbeom
err分享
err收藏
Influence of number of drugs on olfaction in the elderly
err2018-09-01
err0
errOAAI
errG. Ottaviano; E. Savietto; B. Scarpa; A. Bertocco; P. Maculan; G. Sergi; A. Martini; E. Manzato; G. Marioni
err分享
err收藏
学者 查看更多内容