arrow
Return

CoMan: Managing Bandwidth Across Computing Frameworks in Multiplexed Datacenters

delete2018-05-01
delete9
PRE
AI
W
Wenxin Li
D
Deke Guo
A
Alex X. Liu *
K
Keqiu Li
H
Heng Qi *
S
Song Guo
A
Ali Munir
陶小旖 (Xiaoyi Tao)
DOI:10.1109/TPDS.2017.2788003delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Inefficient bandwidth sharing in a datacenter network, between different application frameworks, e.g., MapReduce and Spark, can lead to inelastic and skewed usage of link bandwidth and increased completion times for the applications. Existing work, however, either solely focuses on managing computation and storage resources or controlling only sending/receiving rate at hosts. In this paper, we present CoMan, a solution that provides global in-network bandwidth management in multiplexed data centers, with two goals: improving bandwidth utilization and reducing application completion time. CoMan first designs a novel abstraction of virtual link groups (VLGs) to establish a shared bandwidth resource pool. Based on this pool, CoMan implements a three-level bandwidth allocation model, which enables elastic bandwidth sharing among computing frameworks as well as guarantees network performance for the applications. CoMan further improves the bandwidth utilization by devising a VLG dependency graph and solves an optimization problem to guide the path selection using 3/2-approximation algorithm. We conduct comprehensive trace-driven simulations as well as small-scale testbed experiments to evaluate the performance of CoMan. Extensive simulation results show that CoMan improves the bandwidth utilization and speeds up the application completion time by up to 2.83 x and 6.68 x, respectively, compared to the ECMP + ElasticSwitch solution. Our implementation also verifies that CoMan can realistically speed up the application completion times by 2.32 x on average.
Keywords:
Data-parallel computing frameworks
multiplexed datacenter
bandwidth management
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

H
hong kong polytechnic university
Scholars:
3.0W
Papers: 4.1W
Citations: 921
D
Dalian University of Technology
Scholars:
5.9W
Papers: 4.4W
Citations: 5.5W
N
national university of defense technology - china
Scholars:
1.8W
Papers: 1.4W
Citations: 9
M
michigan state university
Scholars:
3.6W
Papers: 3.2W
Citations: 44
researcher View more organizations