返回
Learning Rate Dropout
DOI:10.1109/TNNLS.2022.3155181.png)
摘要
En 中文
Optimization algorithms are of great importance to efficiently and effectively train a deep neural network. However, the existing optimization algorithms show unsatisfactory convergence behavior, either slowly converging or not seeking to avoid bad local optima. Learning rate dropout (LRD) is a new gradient descent technique to motivate faster convergence and better generalization. LRD aids the optimizer to actively explore in the parameter space by randomly dropping some learning rates (to 0); at each iteration, only parameters whose learning rate is not 0 are updated. Since LRD reduces the number of parameters to be updated for each iteration, the convergence becomes easier. For parameters that are not updated, their gradients are accumulated (e.g., momentum) by the optimizer for the next update. Accumulating multiple gradients at fixed parameter positions gives the optimizer more energy to escape from the saddle point and bad local optima. Experiments show that LRD is surprisingly effective in accelerating training while preventing overfitting.
Keyword:
Training
Convergence
Optimization
Neural networks
Perturbation methods
Informatics
Adaptation models
Neural network
optimization algorithm
regularization
期刊
IF:
8.9
论文数:
7.5K
被引数:
7.2W
机构
引用论文
Distant regulatory elements in a Sox10‐βGEO BAC transgene are required for expression of Sox10 in the enteric nervous system and other neural crest‐derived tissuesSox10-βgeo BAC转基因中的远距离调控元件是肠神经系统和其他神经源性组织中 Sox10 表达所必需的
Synthesis of n-type semiconducting diamond film using diphosphorus pentaoxide as the doping source以五氧化二磷为掺杂源合成n型半导体金刚石膜

