arrow
Return

Enhancing in-memory efficiency for MapReduce-based data processing

delete2018-10-01
delete6
PRE
AI
J
Jorge Veiga *
R
Roberto R. Expósito
G
Guillermo L. Taboada
J
Juan Touriño
DOI:10.1016/j.jpdc.2018.04.001delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
As the memory capacity of computational systems increases, the in-memory data management of Big Data processing frameworks becomes more crucial for performance. This paper analyzes and improves the memory efficiency of Flame-MR, a framework that accelerates Hadoop applications, providing valuable insight into the impact of memory management on performance. By optimizing memory allocation, the garbage collection overheads and execution times have been reduced by up to 85% and 44%, respectively, on a multi-core cluster. Moreover, different data buffer implementations are evaluated, showing that off heap buffers achieve better results overall. Memory resources are also leveraged by caching intermediate results, improving iterative applications by up to 26%. The memory-enhanced version of Flame-MR has been compared with Hadoop and Spark on the Amazon EC2 cloud platform. The experimental results have shown significant performance benefits reducing Hadoop execution times by up to 65%, while providing very competitive results compared to Spark. (C) 2018 Elsevier Inc. All rights reserved.
Keywords:
Big data
MapReduce
In-memory computing
Garbage collector (GC)
Performance evaluation
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

U
Universidade da Coruna
Scholars:
6.6K
Papers: 5.7K
Citations: 11