1
Return

On the Convergence, Implicit Bias, and Edge of Stability of Gradient Descent in Deep Learning: Reviewing recent progress [Special Issue on the Mathematics of Deep Learning]

delete2026-06-12
delete0
PRE
AI
M
Min, Hancheng
L
Lachlan Ewen MacDonald
R
René Vidal
DOI:10.1109/MSP.2026.3665010delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep neural networks (DNNs) trained via gradient descent (GD) with random initialization and without any regularization enjoy good generalization performance in practice despite being highly overparametrized. To theoretically understand this puzzling phenomenon, many works on convergence analysis for GD algorithms on NNs have been developed over the last half-decade. In this article, we review these research efforts and discuss how they address three specific questions related to this puzzle. The first question is why GD finds a global minimum efficiently, which the literature has addressed by studying what level of overparametrization (width, depth, etc.) and what type of initialization lead to a benign optimization landscape along the training trajectory, facilitating a linear convergence rate of GD. The next question is why the global minimum found by GD generalizes well, which has been addressed by showing that overparametrization induces an implicit simplicity bias along the GD trajectory. More recently, it has been observed that in practice, training DNNs with far larger learning rates than theoretically permissible results in faster convergence and better generalization. This leads to the third question of why faster convergence and better generalization can be achieved with a large learning rate, which has been recently addressed by identifying a self-stabilization mechanism and implicit bias toward flat minima.
Keywords:
Artificial neural networks
Training
Gradient methods
Stability analysis
Optimization methods
Sampling methods
Parameter estimation
Signal processing algorithms
Incremental learning
Deep learning

Journal

IEEE Signal Processing Magazine cover
IEEE Signal Processing Magazine
IF:
9.6
Papers:
1.1W
Citations:
1.7W

Organization

S
shanghai jiao tong university
Scholars:
15.1W
Papers: 11.5W
Citations: 159
U
University of Pennsylvania
Scholars:
1.0W
Papers: 3.7K
Citations: 11.8W
Cited Papers

Cited Papers

Citing Papers

Citing Papers