arrow
返回

Spherical perspective on learning with normalization layers

delete2022-05-01
delete3
delete
OA
AI
S
Simon Roburin *
Y
Yann de Mont-Marin
A
Andrei Bursuc
R
Renaud Marlet
P
Patrick Pérez
M
Mathieu Aubry
DOI:10.1016/j.neucom.2022.02.021delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Normalization Layers (NLs) are widely used in modern deep-learning architectures. Despite their apparent simplicity, their effect on optimization is not yet fully understood. This paper introduces a spherical framework to study the optimization of neural networks with NLs from a geometric perspective. Concretely, the radial invariance of groups of parameters, such as filters for convolutional neural networks, allows to translate the optimization steps on the L-2 unit hypersphere. This formulation and the associated geometric interpretation shed new light on the training dynamics. Firstly, the first effective learning rate expression of Adam is derived. Then the demonstration that, in the presence of NLs, performing Stochastic Gradient Descent (SGD) alone is actually equivalent to a variant of Adam constrained to the unit hypersphere, stems from the framework. Finally, this analysis outlines phenomena that previous variants of Adam act on and their importance in the optimization process are experimentally validated. (C) 2022 Elsevier B.V. All rights reserved.
Keyword:
Optimization
Deep learning
Normalization layers
Batch normalization
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

C
centre national de la recherche scientifique (cnrs)
学者数:
24.5W
论文数: 18.2W
被引数: 279
U
universite gustave-eiffel
学者数:
5.6K
论文数: 4.8K
被引数: 5
ESIEE Paris 封面图
ESIEE Paris
学者数:
190
论文数: 145
被引数: 68
学者 查看更多机构
引用论文

引用论文

暂无论文信息