返回
Multiplicative update rules for accelerating deep learning training and increasing robustness
DOI:10.1016/j.neucom.2024.127352.png)
摘要
En 中文
Even nowadays, where Deep Learning (DL) has achieved state-of-the-art performance in a wide range of research domains, accelerating training and building robust DL models remains a challenging task. To this end, generations of researchers have pursued developing robust methods for training DL architectures that can be less sensitive to weight distributions, model architectures, and loss landscapes. However, such methods are limited to adaptive learning rate optimizers, initialization schemes, and clipping gradients without investigating the fundamental rule of parameters update. Although multiplicative updates have contributed significantly to the early development of machine learning and hold strong theoretical claims, to the best of our knowledge, this is the first work that investigates them in a unified manner in the context of DL optimization. In this work, we propose an optimization framework that fits a wide range of optimization algorithms and enables one to apply alternative update rules. To this end, we propose a novel multiplicative update rule and we extend their capabilities by combining it with a traditional additive update term, under a novel hybrid update method. We claim that the proposed framework accelerates training while leading to more robust models in contrast to traditionally used additive update rule, and we experimentally demonstrate its effectiveness in a wide range of task and optimization methods. Such tasks range from convex and non -convex optimization to difficult image classification benchmarks applying a wide range of traditionally used optimization methods and Deep Neural Network (DNN) architectures, providing quantitative and qualitative experimental results.
Keyword:
Robust training framework
Multiplicative update rule
Multiplicative optimizer
Hybrid optimizer
Training acceleration
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
Adaptive multiplicative updates for quadratic nonnegative matrix factorization二次非负矩阵分解的自适应乘法更新
NEUROCOMPUTING
IF6.5
Locally Ordered Graphitized Carbon Coating Enables Recycled Microsized Silicon as High-Performance Anodes局部有序石墨化碳涂层使回收的微米级硅作为高性能负极成为可能

