arrow
返回

Gradient descent optimizes over-parameterized deep ReLU networks

delete2019-10-23
delete232
delete
OA
AI
D
Difan Zou
Y
Yuan Cao
D
Dongruo Zhou
Q
Quanquan Gu *
DOI:10.1007/s10994-019-05839-6delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
We study the problem of training deep fully connected neural networks with Rectified Linear Unit (ReLU) activation function and cross entropy loss function for binary classification using gradient descent. We show that with proper random weight initialization, gradient descent can find the global minima of the training loss for an over-parameterized deep ReLU network, under certain assumption on the training data. The key idea of our proof is that Gaussian random initialization followed by gradient descent produces a sequence of iterates that stay inside a small perturbation region centered at the initial weights, in which the training loss function of the deep ReLU networks enjoys nice local curvature properties that ensure the global convergence of gradient descent. At the core of our proof technique is (1) a milder assumption on the training data; (2) a sharp analysis of the trajectory length for gradient descent; and (3) a finer characterization of the size of the perturbation region. Compared with the concurrent work (Allen-Zhu et al. in A convergence theory for deep learning via over-parameterization, 2018a; Du et al. in Gradient descent finds global minima of deep neural networks, 2018a) along this line, our result relies on milder over-parameterization condition on the neural network width, and enjoys faster global convergence rate of gradient descent for training deep neural networks.
Keyword:
Deep neural networks
Gradient descent
Over-parameterization
Random initialization
Global convergence
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Machine Learning 封面图
Machine Learning
IF:
2.9
论文数:
2.7K
被引数:
3.4W

机构

University of California System 封面图
University of California System
学者数:
37.7W
论文数: 33.8W
被引数: 6.6K
引用论文

引用论文

Charge-Transfer and Energy-Transfer Processes in π-Conjugated Oligomers and Polymers:  A Molecular Picture
err2004-09-28
err0
PREAI
errJean-Luc Brédas; David Beljonne; Veaceslav Coropceanu; Jérôme Cornil
err分享
err收藏
Deep Neural Networks for Acoustic Modeling in Speech Recognition深度神经网络在语音识别声学建模中的应用
err2012-11-01
err8.2K
PREAI
errHinton, Geoffrey; Deng, Li; Yu, Dong; Dahl, George E.; Mohamed, Abdel-rahman; Jaitly, Navdeep; Senior, Andrew; Vanhoucke, Vincent; Patrick Nguyen; Sainath, Tara N.; Kingsbury, Brian
err分享
err收藏
没有更多内容