arrow
Return

Exploiting Stragglers in Distributed Computing Systems With Task Grouping

delete2024-11-01
delete0
PRE
AI
T
Tharindu Adikari *
H
Haider Al-Lawati
J
Jason Lam
胡振华 (Zhenhua Hu)
S
Stark C. Draper
DOI:10.1109/TSC.2024.3495513delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We consider the problem of stragglers in distributed computing systems. Stragglers, which are compute nodes that unpredictably slow down, often increase the completion times of tasks. One common approach to mitigating stragglers is work replication, where only the first completion among replicated tasks is accepted, discarding the others. However, discarding work leads to resource wastage. In this article, we propose a method for exploiting the work completed by stragglers rather than discarding it. The idea is to increase the granularity of the assigned work, and to increase the frequency of worker updates. We show that the proposed method reduces the completion time of tasks via experiments performed on a simulated cluster as well as on Amazon EC2 with Apache Hadoop.
Keywords:
Clustering algorithms
Standards
Prevention and mitigation
Optimization
Hardware
Stochastic processes
Online services
Encyclopedias
Approximation algorithms
Proposals
Distributed systems
stragglers
task scheduling

Journal

IEEE Transactions on Services Computing cover
IEEE Transactions on Services Computing
IF:
5.8
Papers:
2.1K
Citations:
6.5K

Organization

U
university of toronto
Scholars:
14.7W
Papers: 12.0W
Citations: 165