arrow
Return

PAHInA: Precision-Aware Hierarchical In-Network Aggregation for Edge Distributed Training

delete2026-02-03
delete0
PRE
AI
Y
Yingpu Nian
B
Bo Yi
Q
Qiang He
X
X. J. Wang
G
Geyong Min
李克勤 cover
李克勤 (Keqin Li)
S
Sajal K. Das
DOI:10.1109/TON.2026.3660333delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The rise of edge intelligence is driving distributed machine learning toward a new paradigm of edge-collaborative computing. To overcome the severe communication bottleneck in this paradigm, In-Network Aggregation is a critical enabling technology. However, its effectiveness is fundamentally undermined by the profound resource heterogeneity of edge networks. Specifically, edge devices, adapting to hardware constraints, operate at varying numerical precisions, leading to significant data inflation as gradients are aggregated. Compounding this, unevenly distributed network resources and traditional, precision-oblivious routing strategies often misallocate critical, high-precision gradients to low-quality paths. This mismatch creates severe network congestion, crippling the efficiency of distributed training. To address this, we propose the Precision-Aware Hierarchical In-Network Aggregation (PAHInA) framework, the first, to our knowledge, to perform routing optimization for in-network aggregation that explicitly considers precision heterogeneity. The core of PAHInA is an intelligent control-plane scheduler that co-optimizes for gradient priority and path cost, dynamically planning the most cost-effective aggregation strategy for each flow. This fine-grained scheduling guarantees that high-priority gradients are routed through premium, low-latency paths, minimizing global communication overhead. On the data plane, we leverage the eXpress Data Path (XDP) for high-performance packet processing to reduce aggregation-induced overhead. Extensive simulations show that, compared to state-of-the-art baselines, PAHInA significantly mitigates network congestion, reducing end-to-end communication time by up to 33% and boosting overall training throughput by approximately 30%.
Keywords:
Edge computing
distributed training
heterogeneous precision
in-network aggregation
XDP

Journal

I
IEEE Transactions on Networking
IF:
0
Papers:
543
Citations:
0

Organization

S
state university of new york
Scholars:
709
Papers: 451
Citations: 0
U
university of exeter
Scholars:
2.6K
Papers: 1.4K
Citations: 0
N
northeastern university
Scholars:
4.4K
Papers: 1.9K
Citations: 2
M
missouri university of science and technology
Scholars:
467
Papers: 275
Citations: 0
researcher View more organizations