Return
A Survey on Parameter Server Architecture: Approaches for Optimizing Distributed Centralized Learning
DOI:10.1109/ACCESS.2025.3535085.png)
Abstract
En 中文
Deep learning has emerged as a cornerstone technology across various domains, from image classification to natural language processing. However, the computational and data demands of training large-scale neural networks pose significant challenges. Distributed learning approaches, particularly those leveraging data parallelism, have become critical to addressing these challenges. Among these, the parameter server architecture stands out as a widely adopted and scalable solution, enabling efficient training of large models across distributed systems. This survey provides a comprehensive exploration of the parameter server architecture, detailing its design principles and operation. It categorizes and critically analyzes research advancements across five key aspects: consistency control, network optimization, parameter management, straggler handling, and fault tolerance. By synthesizing insights from a wide range of studies, this work highlights the trade-offs and practical effectiveness of various approaches while identifying open challenges and future research directions. The survey aims to serve as a foundational resource for researchers and practitioners striving to enhance the performance and scalability of distributed deep learning systems.
Keywords:
Servers
Training
Data models
Computational modeling
Deep learning
Surveys
Parallel processing
Training data
Solid modeling
Neural networks
distributed systems
network
parameter server
synchronization
centralized
Journal
IF:
3.6
Papers:
9.8W
Citations:
29.4W

