arrow
Return

A Survey on Parameter Server Architecture: Approaches for Optimizing Distributed Centralized Learning

delete2025-01-01
delete0
delete
OA
AI
N
Nikodimos Provatas *
I
Ioannis Konstantinou
N
Nectarios Koziris
DOI:10.1109/ACCESS.2025.3535085delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep learning has emerged as a cornerstone technology across various domains, from image classification to natural language processing. However, the computational and data demands of training large-scale neural networks pose significant challenges. Distributed learning approaches, particularly those leveraging data parallelism, have become critical to addressing these challenges. Among these, the parameter server architecture stands out as a widely adopted and scalable solution, enabling efficient training of large models across distributed systems. This survey provides a comprehensive exploration of the parameter server architecture, detailing its design principles and operation. It categorizes and critically analyzes research advancements across five key aspects: consistency control, network optimization, parameter management, straggler handling, and fault tolerance. By synthesizing insights from a wide range of studies, this work highlights the trade-offs and practical effectiveness of various approaches while identifying open challenges and future research directions. The survey aims to serve as a foundational resource for researchers and practitioners striving to enhance the performance and scalability of distributed deep learning systems.
Keywords:
Servers
Training
Data models
Computational modeling
Deep learning
Surveys
Parallel processing
Training data
Solid modeling
Neural networks
distributed systems
network
parameter server
synchronization
centralized

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

N
National Technical University of Athens
Scholars:
9.7K
Papers: 9.5K
Citations: 8.2K