返回
LipschitzLR: Using theoretically computed adaptive learning rates for fast convergence
DOI:10.1007/s10489-020-01892-0.png)
摘要
En 中文
We present a novel theoretical framework for computing large, adaptive learning rates. Our framework makes minimal assumptions on the activations used and exploits the functional properties of the loss function. Specifically, we show that the inverse of the Lipschitz constant of the loss function is an ideal learning rate. We analytically compute formulas for the Lipschitz constant of several loss functions, and through extensive experimentation, demonstrate the strength of our approach using several architectures and datasets. In addition, we detail the computation of learning rates when other optimizers, namely, SGD with momentum, RMSprop, and Adam, are used. Compared to standard choices of learning rates, our approach converges faster, and yields better results.
Keyword:
Lipschitz constant
Adaptive learning
Machine learning
Deep learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.5
论文数:
7.6K
被引数:
1.7W
机构
引用论文
Facile Synthesis, Characterization, Nanocrystal Growth and Photoluminescence Properties of GeS Nanowires
Nano
IF0
Mechanical, but not infective, pacemaker erosion may be successfully managed by re-implantation of pacemakers.
Heart
IF0
Structural study of lanthanides(III) in aqueous nitrate and chloride solutions by EXAFS通过EXAFS对硝酸盐和氯化物水溶液中镧系元素 (III) 的结构研究
Accurate quantitative estimation of energy performance of residential buildings using statistical machine learning tools
ENERGY AND BUILDINGS
IF7.1

