返回
Softmax-kernel reproduced gradient descent for stochastic optimization on data
DOI:10.1016/j.sigpro.2025.109904.png)
摘要
En 中文
Stochastic gradient descent (SGD) is commonly used for machine learning on streaming data. However, it suffers from slow convergence due to gradient variance. To address this issue, the Reproducing Kernel Hilbert Space (RKHS) theory is applied to build a kernel learning model and learn the gradient of the risk function. To avoid the inherent dimensional trap in kernel methods, a softmax kernel function is designed to reproduce the gradient iteratively, by which a novel algorithm called softmax-kernel reproduced gradient descent (SoKRGD) is further proposed. It is shown that SoKRGD achieves a faster convergence rate than SGD. Experimental results are provided to validate these findings by training ResNet50 and Vision Transformer (ViT). It is observed that using the reproduced gradient in place of the stochastic gradient can promote the performance of SGD-based optimizers.
Keyword:
Streaming data
Stochastic optimization
Reproducing Kernel Hilbert Space
Stochastic gradient descent
Kernel function
期刊
IF:
3.6
论文数:
9.9K
被引数:
1.7W

