arrow
返回

Two-Step Nyström Sampling for Large-Scale Kernel Approximation

delete2025-10-07
delete0
PRE
AI
L
Li He
H
Hong Zhang
DOI:10.1109/TBDATA.2025.3618472delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Nyström approximation is one of the most popular approximation methods to accelerate kernel analysis on large-scale data sets. Nyström employs one single landmark set to obtain eigenvectors (low-rank decomposition) and projects the entire data set to the eigenvectors (embedding). Most existing methods focus on accelerating landmark selection. For extremely large-scale data sets, however, the embedding time cost, rather than that of low-rank decomposition, is critical. In addition, both accuracy and embedding time cost are dominated by the landmark set size. As a result, using more landmarks is the only way to improve accuracy at the cost of extremely high embedding costs. In this paper, we propose a method for the first time to decouple embedding cost from that of low-rank decomposition. We first obtain the eigenvectors from a large landmark set for a low error, and then optimize a small landmark set that minimizes the landmark-set-embedding error to ensure a low embedding cost. In return, our accuracy is close to that of the large landmark set but the small one dominates the embedding time cost. Our method can deal with popular kernels and be plugged into most existing methods. Experimental results demonstrate the superiority of the proposed method.
Keyword:
Clustering
kernel approximation
large-scale data
nyström approximation

期刊

I
IEEE Transactions on Big Data
IF:
5.7
论文数:
860
被引数:
3.0K

机构

S
southern university of science and technology
学者数:
4.7K
论文数: 1.7K
被引数: 0
引用论文

引用论文

Fast Communication-Efficient Spectral Clustering over Distributed Data
err2021-03-01
err2
errOAAI
errYan, Donghui; Wang, Yingjie; Wang, Jin; Wu, Guodong; Wang, Honggang
err分享
err收藏
Nyströmformer: A Nyström-based Algorithm for Approximating Self-Attention
err2021-05-18
err0
errOAAI
errYunyang Xiong; Zhanpeng Zeng; Rudrasis Chakraborty; Mingxing Tan; Glenn Fung; Yin Li; Vikas Singh
err分享
err收藏
Fast Large-Scale Spectral Clustering via Explicit Feature Mapping
err2019-03-01
err79
PREAI
errHe, Li; Ray, Nilanjan; Guan, Yisheng; Zhang, Hong
err分享
err收藏
学者 查看更多内容