返回
HAGC: A Hardware-Aware Gradient Compression framework for distributed deep learning
DOI:10.1016/j.sysarc.2026.103770.png)
摘要
En 中文
• 双边Hadamard变换将工作负载转移到张量核心。
• 错误反馈和协同设计确保吞吐量和收敛稳定性。
• 在A100 GPU上实现了最高3.15倍的加速和2.9倍的能耗降低。
Keyword:
Hadamard Transform
Gradient Compression
Tensor Cores
Distributed Deep Learning
Energy Efficiency
期刊
IF:
4.1
论文数:
3.0K
被引数:
4.2K
机构
引用论文
暂无论文信息

