arrow
Return

Communication-efficient ADMM-based distributed algorithms for sparse training

delete2023-09-01
delete3
PRE
AI
G
Guozheng Wang
Y
Yongmei Lei *
Y
Yongwen Qiu
L
Lingfei Lou
Y
Yixin Li
DOI:10.1016/j.neucom.2023.126456delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In large-scale distributed machine learning (DML), the synchronization efficiency of the distributed algorithm becomes a critical factor that affects the training time of machine learning models as the computing scale increases. To address this challenge, we propose a novel algorithm called Grouped Sparse AllReduce based on the 2D-Torus topology (2D-TGSA), which enables constant transmission traffic that does not change with the number of workers. Our experimental results demonstrate that 2D-TGSA outperforms several benchmark algorithms in terms of synchronization efficiency. Moreover, we integrate the general form consistent ADMM with 2D-TGSA to develop a distributed algorithm (2D-TGSAADMM) that exhibits excellent scalability and can effectively handle large-scale distributed optimization problems. Furthermore, we enhance 2D-TGSA-ADMM by adopting the resilient adaptive penalty parameter approach, resulting in a new algorithm called 2D-TGSA-TPADMM. Our experiments on training the logistic regression model with '1-norm on the Tianhe-2 supercomputing platform demonstrate that our proposed algorithm can significantly reduce the synchronization time and training time compared to state-of-the-art methods.& COPY; 2023 Elsevier B.V. All rights reserved.
Keywords:
ADMM
Grouped Sparse AllReduce
Two-dimensional torus topology
Synchronization algorithm

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

S
shanghai university
Scholars:
3.9W
Papers: 2.7W
Citations: 52
Cited Papers

Cited Papers

A Self-Driving Lab for Nano- and Advanced Materials Synthesis
err
IF0
err2024-12-09
err0
errOAAI
errMohammad Zaki; Carsten Prinz; Bastian Ruehle
errShare
errSave
Machine Learning for Molecular Simulation
err2020-04-20
err562
errOAAI
errNoe, Frank; Tkatchenko, Alexandre; Mueller, Klaus-Robert; Clementi, Cecilia
errShare
errSave
A Survey on Distributed Machine Learning
err2020-03-20
err430
errOAAI
errVerbraeken, Joost; Wolting, Matthijs; Katzy, Jonathan; Kloppenburg, Jeroen; Verbelen, Tim; Rellermeyer, Jan S.
errShare
errSave
Machine learning in medical applications: A review of state-of-the-art methods
err2022-06-01
err220
PREAI
errShehab, Mohammad; Abualigah, Laith; Shambour, Qusai; Abu-Hashem, Muhannad A.; Shambour, Mohd Khaled Yousef; Alsalibi, Ahmed Izzat; Gandomi, Amir H.
errShare
errSave
TensorOpt: Exploring the Tradeoffs in Distributed DNN Training With Auto-Parallelism
err2022-08-01
err17
errOAAI
errCai, Zhenkun; Yan, Xiao; Ma, Kaihao; Wu, Yidi; Huang, Yuzhen; Cheng, James; Su, Teng; Yu, Fan
errShare
errSave
errShare
errSave
researcher View more