arrow
Return

Distributed algorithm for best subset regression

delete2025-06-01
delete0
PRE
AI
H
Hao Ming
DOI:10.1016/j.eswa.2025.127224delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
High-dimensional massive data modeling faces critical challenges in computational efficiency, memory constraints, and privacy protection. We develop a distributed framework for best subset regression with convex twice-differentiable losses (e.g., linear, multiplicative, and logistic regression). The proposed distributed enhanced primal-dual active set (DEPDAS) algorithm employs enhanced distributed computing to efficiently approximate optimal solutions in low-dimensional parameter spaces. Under standard regularity conditions, DEPDAS preserves the statistical properties of the full-sample-based EPDAS algorithm, including optimal estimation error rates and Oracle properties. With a per-iteration communication cost of O(2T+2p) for DEPDAS, our master-machine initialization strategy accelerates convergence while reducing communication overhead. Furthermore, we derive a lower communication DEPDAS (LCDEPDAS) variant with O(4T) per-iteration cost. Extensive simulations and empirical studies demonstrate the superiority of both algorithms over state-of-the-art methods in estimation accuracy and prediction performance.
Keywords:
Distributed learning
L0 penalty
KKT conditions
Oracle property

Journal

Expert Systems with Applications cover
Expert Systems with Applications
IF:
7.5
Papers:
2.9W
Citations:
10.2W

Organization

No organization information available