arrow
返回

Powered stochastic optimization with hypergradient descent for large-scale learning systems

delete2024-03-01
delete0
PRE
AI
杨壮 封面图
杨壮 (Zhuang Yang) *
X
X. L. Li
DOI:10.1016/j.eswa.2023.122017delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Stochastic optimization (SO) algorithms based on the Powerball function, namely powered stochastic optimization (PoweredSO) algorithms, have been confirmed, effectively, and demonstrated great potential in the context of large-scale optimization and machine learning tasks. Nevertheless, the issue of how to determine the learning rate for PoweredSO is a challenge and still unsolved problem. In this paper, we propose a class of adaptive PoweredSO approaches that are efficient, scalable and robust. It takes advantage of the hypergradient descent (HD) technique to automatically acquire an online learning rate for PoweredSO-like methods. In the first part, we study the behavior of the canonical PoweredSO algorithm, the Powerball stochastic gradient descent (pbSGD) method, with HD. The existing PoweredSO algorithms also suffer from the high variance because they take the similar algorithmic framework to SO algorithms, arising from sampling tactics. Therefore, the second portion develops an adaptive powered variance-reduced optimization method via utilizing both variance-reduced technique and HD. Moreover, we present the convergence analysis of the proposed algorithms and explore their iteration complexity on non-convex cases. Numerical experiments are conducted on machine learning tasks, verifying the superior performance over modern SO algorithms.
Keyword:
Powerball function
Stochastic optimization
Variance reduction
Hypergradient descent
Adaptive learning rate

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
3.0W
被引数:
10.2W

机构

S
soochow university - china
学者数:
5.2W
论文数: 3.6W
被引数: 82
引用论文

引用论文

The influence of workload levels on performance in a rural hospital
err2015-12-02
err0
PREAI
errJames Avoka Asamani; Ninon P Amertil; Margaret Chebere
err分享
err收藏
err分享
err收藏
err分享
err收藏
k-best feature selection and ranking via stochastic approximation基于随机逼近的k-best特征选择和排序
err2023-03-01
err10
PREAI
errAkman, David V.; Malekipirbazari, Milad; Yenice, Zeren D.; Yeo, Anders; Adhikari, Niranjan; Wong, Yong Kai; Abbasi, Babak; Gumus, Alev Taskin
err分享
err收藏
Accelerated stochastic gradient descent with step size selection rules
err2019-06-01
err20
PREAI
errYang, Zhuang; Wang, Cheng; Zhang, Zhemin; Li, Jonathan
err分享
err收藏
Neural Importance Sampling
err2019-10-10
err180
errOAAI
errMueller, Thomas; Mcwilliams, Brian; Rousselle, Fabrice; Gross, Markus; Novak, Jan
err分享
err收藏
Mini-batch algorithms with online step size具有在线步长的小批量算法
err2019-02-01
err25
PREAI
errYang, Zhuang; Wang, Cheng; Zhang, Zhemin; Li, Jonathan
err分享
err收藏
学者 查看更多内容