Return
The Optimal Condition Number for ReLU Function
DOI:10.1109/TIT.2025.3633042.png)
Abstract
En 中文
ReLU is a widely used activation function in deep neural networks. This paper explores the stability properties of the ReLU map. For any weight matrix A is an element of R-mxn and bias vector b is an element of R-m at a given layer, we define the condition number kappa(A,b) as kappa(A,b) = U-A,U-b / L-A,L-b , where U-A,U-b and L-A,L-b are the upper and lower Lipschitz constants, respectively. We first demonstrate that for any given A and b, the condition number satisfies root kappa(A,b) >= root 2. Moreover, when the weights of the network at a given layer are initialized as random i.i.d. Gaussian variables and the bias term is set to zero, the condition number asymptotically approaches this lower bound. Our findings offer valuable insights into the characteristics of randomly initialized neural networks, contributing to a better understanding of their initial behavior and potential performance.
Keywords:
Lower bound
Vectors
Standards
Nonlinear distortion
Upper bound
Linear matrix inequalities
Biological neural networks
Asymptotic stability
Artificial neural networks
Thermal stability
ReLU
condition number
bi-Lipschitz constants
random initialization
Journal
I
IF:
2.9
Papers:
317
Citations:
0

