arrow
Return

Distributed Computing and Inference for Big Data

delete2024-04-22
delete0
delete
OA
AI
L
Ling Zhou *
Z
Ziyang Gong
P
Pengcheng Xiang
DOI:10.1146/annurev-statistics-040522-021241delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Data are distributed across different sites due to computing facility limitations or data privacy considerations. Conventional centralized methods-those in which all datasets are stored and processed in a central computing facility-are not applicable in practice. Therefore, it has become necessary to develop distributed learning approaches that have good inference or predictive accuracy while remaining free of individual data or obeying policies and regulations to protect privacy. In this article, we introduce the basic idea of distributed learning and conduct a selected review on various distributed learning methods, which are categorized by their statistical accuracy, computational efficiency, heterogeneity, and privacy. This categorization can help evaluate newly proposed methods from different aspects. Moreover, we provide up-to-date descriptions of the existing theoretical results that cover statistical equivalency and computational efficiency under different statistical learning frameworks. Finally, we provide existing software implementations and benchmark datasets, and we discuss future research opportunities.
Keywords:
communication efficiency
distributed learning
federated learning
heterogeneity
statistical equivalence

Journal

Annual Review of Statistics and Its Application cover
Annual Review of Statistics and Its Application
IF:
8.7
Papers:
211
Citations:
2.4K

Organization

S
southwestern university of finance & economics - china
Scholars:
3.0K
Papers: 3.4K
Citations: 4