arrow
Return

An Efficient Computing and Communication Framework for Large-Scale Data Processing Cluster

delete2026-07-09
delete0
PRE
AI
X
Xuya Jia
Z
Zhiyi Yao
C
Chao Peng
Z
Zihao Zhao
E
Edison Liu
C
Congcong Miao
徐
徐跃东 (Yuedong Xu)
DOI:10.1109/ton.2026.3711403delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Data processing clusters dealing with big data are facing extended completion times for jobs because the RDMA feature is not being used efficiently. Our production data from a large cluster, which has many server nodes and is responsible for processing extensive data jobs, reveals that the current use of the RDMA technology is causing some jobs to finish much later than usual, with a few taking more than double the standard time to be completed. In this paper, we introduce the design and implementation of Turbo, a high-performance, scalable communication framework tailored for large-scale data processing clusters. The essence of Turbo’s strategy lies in the use of a dynamic block-level flowlet transmission system and a non-blocking communication middleware, which are designed to boost network throughput and system scalability. Moreover, Turbo maintains high system reliability by incorporating an external shuffle service with TCP as a fail-safe option and maintaining metadata management information using the NICs. We have integrated Turbo into Apache Spark and conducted evaluations on both a small-scale test environment and a large-scale cluster with hundreds of server nodes. The findings from the small-scale testbed demonstrate that Turbo enhances network throughput by 15.1% and upholds high system reliability. Additionally, the large-scale production data indicates that Turbo is capable of decreasing job completion times by 23.9% and increasing the job completion rate by $2.03\times $ compared to the current RDMA solutions. In addition, during large-scale tests, we also found that Turbo has improved the computing efficiency of the cluster and saved approximately 24.3% of the CPU utilization.
Keywords:
Load balancing
spark RDMA
data processing
communication acceleration

Journal

I
IEEE-ACM Transactions on Networking
IF:
3.6
Papers:
4.4K
Citations:
9.5K

Organization

T
tencent technology company ltd.
Scholars:
3
Papers: 2
Citations: 0
F
fudan university
Scholars:
11.8W
Papers: 7.7W
Citations: 121
N
NVIDIA
Scholars:
90
Papers: 43
Citations: 10
researcher View more organizations
Cited Papers

Cited Papers

errShare
errSave
Congestion Control for Large-Scale RDMA Deployments
err2015-08-17
err0
PREAI
errYibo Zhu; Haggai Eran; Daniel Firestone; Chuanxiong Guo; Marina Lipshteyn; Yehonatan Liron; Jitendra Padhye; Shachar Raindel; Mohamad Haj Yahia; Ming Zhang
errShare
errSave
A Cloud-Optimized Transport Protocol for Elastic and Scalable HPC
err2020-11-01
err0
PREAI
errLeah Shalev; Hani Ayoub; Nafea Bshara; Erez Sabbag
errShare
errSave
errShare
errSave
Pregel
err2010-06-06
err0
PREAI
errGrzegorz Malewicz; Matthew H. Austern; Aart J.C Bik; James C. Dehnert; Ilan Horn; Naty Leiser; Grzegorz Czajkowski
errShare
errSave
PLB
err2022-08-22
err0
errOAAI
errMubashir Adnan Qureshi; Yuchung Cheng; Qianwen Yin; Qiaobin Fu; Gautam Kumar; Masoud Moshref; Junhua Yan; Van Jacobson; David Wetherall; Abdul Kabbani
errShare
errSave
Accelerating Spark with RDMA for Big Data Processing: Early Experiences
err2014-08-01
err0
PREAI
errXiaoyi Lu; Md. Wasi Ur Rahman; Nusrat Islam; Dipti Shankar; Dhabaleswar K. Panda
errShare
errSave
Network Load Balancing with In-network Reordering Support for RDMA
err2023-09-01
err0
errOAAI
errCha Hwan Song; Xin Zhe Khooi; Raj Joshi; Inho Choi; Jialin Li; Mun Choon Chan
errShare
errSave
researcher View more