arrow
Return

D-DOSA:DPU-Based Dataflow Offloading and Sparse Allreduce Framework for Distributed Training

delete2025-11-19
delete0
PRE
AI
Z
Zhenqi Yu
W
Wenjing Li
郭少勇 (Shaoyong Guo)
李庆锋 cover
李庆锋 (Qingfeng Li)
齐峰 cover
齐峰 (Feng Qi)
J
Jiapeng Xiu
DOI:10.1109/TCC.2025.3634564delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Communication overhead represents a primary bottleneck in distributed deep learning, impeding training scalability. Although existing gradient sparsification techniques reduce network traffic, they introduce critical limitations: they fail to optimize intra-node data paths and are incompatible with efficient, decentralized Allreduce operations. To address these issues, we propose D-DOSA, a DPU-based communication offloading framework. D-DOSA incorporates two key innovations: 1) D-DO, an architecture that establishes a direct GPU-DPU data path to offload data loading and intra-node communication from the host CPU; and 2) D-SA, a novel sparse Allreduce algorithm that, for the first time, enables compatibility between sparse tensors and high-performance, ring-based communication. We evaluated D-DOSA on a 8-node, DPU-enabled cluster using representative models including VGG, LSTM, and BERT. Experimental results demonstrate that our framework accelerates training by up to 1.32x compared to the state-of-the-art sparse training baseline, without compromising accuracy. Ultimately, D-DOSA shows that co-designing data-flow architectures and communication algorithms on the DPU resolves key bottlenecks in sparse training and presents a viable path toward scalable performance in larger systems.
Keywords:
Distributed deep learning
data parallelism
allreduce
gradient sparsification
data processing unit

Journal

I
IEEE Transactions on Cloud Computing
IF:
5
Papers:
1.8K
Citations:
4.3K

Organization

B
beijing university of posts and telecommunications
Scholars:
2.0K
Papers: 756
Citations: 0