返回
Robust kernel-based gradient descent with random features
DOI:10.1007/s10444-025-10263-7.png)
摘要
En 中文
In large-scale machine learning, the computational cost of kernel methods can become prohibitive due to the need to compute pairwise kernel evaluations on extensive datasets. The random feature method is one of the most popular techniques for accelerating kernel methods in large-scale problems while maintaining statistical accuracy. In this paper, we investigate the generalization properties of a robust gradient descent algorithm utilizing random features within a statistical learning framework, where we employ the robust loss function l sigma\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$l_{\sigma }$$\end{document} instead of the traditional squared loss during training. This loss function is defined by a windowing function G and a scale parameter sigma\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\sigma $$\end{document}, allowing it to encompass a wide range of commonly used robust losses for regression when G and sigma\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\sigma $$\end{document} are appropriately selected. However, it remains unclear whether the random feature method can preserve statistical accuracy in this context. We analyze the generalization error of the estimator produced by the gradient descent algorithm with random features. Our findings demonstrate that with a suitably chosen scale parameter sigma\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\sigma $$\end{document} and an appropriate number of random features M, our estimator can converge to the regression function in L2\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$L<^>2$$\end{document}-norm at optimal rates in the mini-max sense (up to a logarithmic term), even if the regression function may not reside in the reproducing kernel Hilbert space.
Keyword:
Learning theory
Gradient descent
Robust regression
Random features
Reproducing kernel Hilbert space
期刊
A
IF:
2.1
论文数:
56
被引数:
0

