arrow
Return

Aggregation-based partitioning algorithm for traffic congestion in MapReduce

delete2025-03-25
delete0
delete
OA
AI
S
Shabbir, Aisha
M
Muhammad Imran Hamid
A
Ahmed A. Abd El‐Latif
M
May Almousa *
R
Rania A. Elsayed
W
Waleed Ghaznavi
S
Samia Allaoua Chelloug
A
Abdelhamied A. Ateya *
DOI:10.1186/s40537-025-01115-8delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The current era has witnessed a remarkable transformation in scientific frontiers, largely driven by advancements in the digital domain. This has resulted in an unprecedented explosion of data known as big data. Among the platforms capable of effectively handling massive data volumes cost-effectively, MapReduce stands out. While previous research has focused on enhancing MapReduce's overall performance by selecting and scheduling mappers, little focus has been given to optimizing the shuffle phase's impact on performance. MapReduce operates through multiple phases, with the shuffle phase generating substantial data traffic in the case of heavy jobs. Optimizing or encapsulating aspects, e.g., catalyst, can significantly accelerate the platform's performance. This work introduces an aggregation-based partitioning algorithm (ABPA) that addresses the limitations of existing approaches commonly adopted to reduce traffic congestion during the intermediate, i.e., shuffle, phase of MapReduce. The proposed ABPA algorithm is evaluated through several experiments involving different data sets of varying types and lengths, employing different numbers of partitions, each containing a specified number of mappers and an aggregator. The experimental results demonstrated a significant reduction in network traffic costs and improved total MapReduce job execution time when using the ABPA scheme. Specifically, the proposed algorithm achieved 73% and 56% network traffic cost improvement over basic hash and conventional aggregation schemes, respectively. Additionally, it reduced total job completion time by 60% and 46% compared to the same schemes.
Keywords:
Aggregation-based partitioning
Big data
MapReduce
Job completion time
Shuffle phase optimization
Network traffic cost

Journal

Journal of Big Data cover
Journal of Big Data
IF:
6.4
Papers:
1.4K
Citations:
1.1W

Organization

S
sir syed case inst technol
Scholars:
2
Papers: 2
Citations: 0
M
Menoufia Univ
Scholars:
253
Papers: 166
Citations: 48
N
Natl Univ Sci and Technol
Scholars:
255
Papers: 211
Citations: 79
P
Princess Nourah Bint Abdulrahman Univ
Scholars:
360
Papers: 417
Citations: 103
G
Govt Coll Women Univ
Scholars:
23
Papers: 20
Citations: 6
Z
Zagazig University
Scholars:
5.9K
Papers: 5.0K
Citations: 91
researcher View more organizations