arrow
Return

Adding data analytics capabilities to scaled-out object store

delete2016-11-01
delete0
PRE
AI
C
Cengiz Karakoyunlu *
J
J. Chandy
A
Alma Riska
DOI:10.1016/j.jss.2016.07.029delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This work focuses on enabling effective data analytics on scaled-out object storage systems. Typically, applications perform MapReduce computations by first copying large amounts of data to a separate compute cluster (i.e. a Hadoop cluster). However; this approach is not very efficient considering that storage systems can host hundreds of petabytes of data. Network bandwidth can be easily saturated and the overall energy consumption would increase during large-scale data transfer. Instead of moving data between remote clusters; we propose the implementation of a data analytics layer on an object-based storage cluster to perform in-place MapReduce computation on existing data. The analytics layer is tied to the underlying object store, utilizing its data redundancy and distribution policies across the cluster. We implemented this approach with Ceph object storage system and Hadoop, and conducted evaluations with various benchmarks. Performance evaluations show that initial data copy performance is improved by up to 96% and the MapReduce performance is improved by up to 20% compared to the stock Hadoop implementation. (C) 2016 Elsevier Inc. All rights reserved.
Keywords:
In-situ data analytics
Object storage
Attribute-based storage
MapReduce
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Systems and Software cover
Journal of Systems and Software
IF:
4.1
Papers:
5.4K
Citations:
8.4K

Organization

U
University of Connecticut
Scholars:
2.4W
Papers: 2.2W
Citations: 2.5W
N
netapp, inc.
Scholars:
14
Papers: 13
Citations: 0