arrow
返回

Regularized Top-k: A Bayesian Framework for Gradient Sparsification

delete2025-01-01
delete0
PRE
AI
A
Ali Bereyhi *
B
Ben Liang
G
Gary Boudreau
A
Ali Afana
DOI:10.1109/TSP.2025.3624791delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Error accumulation is effective for gradient sparsification in distributed settings: initially-unselected gradient entries are eventually selected as their accumulated error exceeds a certain level. The accumulation essentially behaves as a scaling of the learning rate for the selected entries. Although this property prevents the slow-down of lateral movements in distributed gradient descent, it can deteriorate convergence in some settings. This work proposes a novel sparsification scheme that controls the learning rate scaling of error accumulation. The development of this scheme follows two major steps: first, gradient sparsification is formulated as an inverse probability (inference) problem, and the Bayesian optimal sparsification mask is derived as a maximum-a-posteriori estimator. Using the prior distribution inherited from Top-k, we derive a new sparsification algorithm which can be interpreted as a regularized form of Top-k. We call this algorithm regularized Top-k (RegTop-k). It utilizes past aggregated gradients to evaluate posterior statistics of the next aggregation. It then prioritizes the local accumulated gradient entries based on these posterior statistics. We validate our derivation through various numerical experiments. In distributed linear regression, it is observed that while Top-k remains at a fixed distance from the global optimum, RegTop-k converges to the global optimum at significantly higher compression ratios. We further demonstrate the generalization of this observation by employing RegTop-k in distributed training of ResNet-18 on CIFAR-10, as well as fine-tuning of multiple computer vision models on the ImageNette dataset. Our numerical results confirm that as the compression ratio increases, RegTop-k sparsification noticeably outperforms Top-k.
Keyword:
Training
Convergence
Signal processing algorithms
Bayes methods
Toy manufacturing industry
Servers
Model compression
Quantization (signal)
Inference algorithms
Distance learning
Gradient sparsification
model compression
distributed learning
neural networks
Bayesian inference

期刊

I
IEEE Transactions on Signal Processing
IF:
5.8
论文数:
318
被引数:
0

机构

U
university of toronto
学者数:
14.8W
论文数: 12.0W
被引数: 165
引用论文

引用论文

Nonparametric Statistical Methods
err2015-07-24
err0
PREAI
errMyles Hollander; Douglas A. Wolfe; Eric Chicken
err分享
err收藏
vqSGD: Vector Quantized Stochastic Gradient Descent
err2022-07-01
err0
errOAAI
errVenkata Gandikota; Daniel Kane; Raj Kumar Maity; Arya Mazumdar
err分享
err收藏
AdaComp : Adaptive Residual Gradient Compression for Data-Parallel Distributed Training
err2018-04-29
err0
errOAAI
errChia-Yu Chen; Jungwook Choi; Daniel Brand; Ankur Agrawal; Wei Zhang; Kailash Gopalakrishnan
err分享
err收藏
A Survey on Distributed Machine Learning分布式机器学习综述
err2020-03-20
err430
errOAAI
errVerbraeken, Joost; Wolting, Matthijs; Katzy, Jonathan; Kloppenburg, Jeroen; Verbelen, Tim; Rellermeyer, Jan S.
err分享
err收藏
学者 查看更多内容