返回
RIEMANNIAN NATURAL GRADIENT METHODS
DOI:10.1137/22M1509643.png)
摘要
En 中文
This paper studies large-scale optimization problems on Riemannian manifolds whose objective function is a finite sum of negative log-probability losses. Such problems arise in various machine learning and signal processing applications. By introducing the notion of Fisher information matrix in the manifold setting, we propose a novel Riemannian natural gradient method, which can be viewed as a natural extension of the natural gradient method from the Euclidean setting to the manifold setting. We establish the almost-sure global convergence of our proposed method under standard assumptions. Moreover, we show that if the loss function satisfies certain convexity and smoothness conditions and the input-output map satisfies a Riemannian Jacobian stability condition, then our proposed method enjoys a local linear---or, under the Lipschitz continuity of the Riemannian Jacobian of the input-output map, even quadratic---rate of convergence. We then prove that the Riemannian Jacobian stability condition will be satisfied by a two-layer fully connected neural network with batch normalization with high probability, provided that the width of the network is sufficiently large. This demonstrates the practical relevance of our convergence rate result. Numerical experiments on applications arising from machine learning demonstrate the advantages of the proposed method over state-of-the-art ones.
Keyword:
manifold optimization
Riemannian Fisher information matrix
Kronecker-factored approximation
natural gradient method
期刊
IF:
2.6
论文数:
5.1K
被引数:
1.8W
机构
引用论文
Effects on the evolution of microstructure and properties of low-silicon cast aluminum alloys in service服役条件下低硅铸铝合 金微观结构与性能的演化效应
A Computational Framework for the Indirect Estimation of Interface Thermal Resistance of Composite Materials Using Xpinns基于Xpinns的复合材料界面热阻间接估计的计算框架

