arrow
Return

aMacP: An adaptive optimization algorithm for Deep Neural Network

delete2025-03-01
delete0
PRE
AI
S
Shubhankar Bhakta
U
Utpal Nandi *
C
Chiranjit Changdar
B
Bachchu Paul
T
Tapas Si
R
Rajat Kumar Pal
DOI:10.1016/j.neucom.2024.129242delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Stochastic gradient-based optimizers are used to train convolutional neural networks (CNNs). Due to its adaptable momentum, the Adam optimizer has recently gained a lot of attention because it addresses the SGD's dying gradient problem but ignores local gradient changes. In order to manage step size, angularGrad calculates a score based just on gradient angular values from prior iterations, and the diffGrad optimizer adjusts step size according to the distinction between the gradients in the current and the recent past. Existing optimizers, however, are still unable to effectively leverage optimization curve information, take advantage of trends in parameter updating to generate adaptive terms, or account for the 1st and 2nd moment for parameter updates. The novel aMacP optimizer proposed in this research takes into account the average of both momentums and also the average of two consecutive parameters for adaptively changing the step size. It is the first effort in the domain to use the average of two momentum variables (the exponential decay average (EDA) of past gradients and the EDA of prior squared gradients) as the denominator for parameter updating to enhance the optimization process. With a more precise step size, the optimization steps become smoother. Yet, rigorous tests comparing state-of-the-art approaches to benchmark datasets show that aMacP performs better. Additionally, it is demonstrated that aMacP performs consistently well when training CNN with various activation functions. Using the Rosenbrock function and Rastrigin function to assess the converged path, these appeared that the suggested aMacP optimizer provides superior results due to its assortment of frequently employed optimization techniques. Additionally, the suggested approach produces improved accuracy fora range of NLP applications and mAP for object detection.
Keywords:
Adam
DiffGrad
AngularGrad
Adaptive learning
Optimization
Neural networks
Average of momentums
Average of parameters

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

V
Vidyasagar University
Scholars:
1.1K
Papers: 967
Citations: 2
U
University of Calcutta
Scholars:
4.5K
Papers: 4.2K
Citations: 3.6K