arrow
返回

Robust kernel-based gradient descent with random features

delete2025-10-28
delete0
PRE
AI
Q
Qi Hong
Z
Zheng-Chu Guo *
DOI:10.1007/s10444-025-10263-7delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In large-scale machine learning, the computational cost of kernel methods can become prohibitive due to the need to compute pairwise kernel evaluations on extensive datasets. The random feature method is one of the most popular techniques for accelerating kernel methods in large-scale problems while maintaining statistical accuracy. In this paper, we investigate the generalization properties of a robust gradient descent algorithm utilizing random features within a statistical learning framework, where we employ the robust loss function l sigma\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$l_{\sigma }$$\end{document} instead of the traditional squared loss during training. This loss function is defined by a windowing function G and a scale parameter sigma\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\sigma $$\end{document}, allowing it to encompass a wide range of commonly used robust losses for regression when G and sigma\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\sigma $$\end{document} are appropriately selected. However, it remains unclear whether the random feature method can preserve statistical accuracy in this context. We analyze the generalization error of the estimator produced by the gradient descent algorithm with random features. Our findings demonstrate that with a suitably chosen scale parameter sigma\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\sigma $$\end{document} and an appropriate number of random features M, our estimator can converge to the regression function in L2\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$L<^>2$$\end{document}-norm at optimal rates in the mini-max sense (up to a logarithmic term), even if the regression function may not reside in the reproducing kernel Hilbert space.
Keyword:
Learning theory
Gradient descent
Robust regression
Random features
Reproducing kernel Hilbert space

期刊

A
Advances in Computational Mathematics
IF:
2.1
论文数:
56
被引数:
0

机构

Z
Zhejiang University
学者数:
1.5W
论文数: 5.2K
被引数: 17.8W
引用论文

引用论文

Learning theory of distributed spectral algorithms
err2017-06-21
err0
PREAI
errZheng-Chu Guo; Shao-Bo Lin; Ding-Xuan Zhou
err分享
err收藏
err分享
err收藏
err分享
err收藏
Learning Theory
err
IF0
err2010-03-05
err0
PREAI
errFelipe Cucker; Ding Xuan Zhou
err分享
err收藏
err分享
err收藏
学者 查看更多内容