arrow
返回

Bidirectionally self-normalizing neural networks

delete2023-10-01
delete4
delete
OA
AI
Y
Yao Lu *
S
Stephen Jay Gould
T
Thalaiyasingam Ajanthan
DOI:10.1016/j.neunet.2023.08.017delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
The problem of vanishing and exploding gradients has been a long-standing obstacle that hinders the effective training of neural networks. Despite various tricks and techniques that have been employed to alleviate the problem in practice, there still lacks satisfactory theories or provable solutions. In this paper, we address the problem from the perspective of high-dimensional probability theory. We provide a rigorous result that shows, under mild conditions, how the vanishing/exploding gradients problem disappears with high probability if the neural networks have sufficient width. Our main idea is to constrain both forward and backward signal propagation in a nonlinear neural network through a new class of activation functions, namely Gaussian-Poincare normalized functions, and orthogonal weight matrices. Experiments on both synthetic and real-world data validate our theory and confirm its effectiveness on very deep neural networks when applied in practice.(c) 2023 Elsevier Ltd. All rights reserved.
Keyword:
Neural networks
Vanishing/exploding gradient problem
Training
Optimization
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Neural Networks 封面图
Neural Networks
IF:
6.3
论文数:
7.9K
被引数:
3.0W

机构

A
Australian National University
学者数:
2.1W
论文数: 2.3W
被引数: 3.9W
引用论文

引用论文

Analysis of Trace VX in Acidified VX Hydrolysate Samples
err
IF0
err2009-07-01
err0
PREAI
errDennis K. Rohrbaugh; Yu-Chu Yang
err分享
err收藏
Resistance Training and Youth阻力训练和青年
err1989-11-01
err0
errOAAI
errWilliam J. Kraemer; Andrew C. Fry; Peter N. Frykman; Brian Conroy; Jay Hoffman
err分享
err收藏