arrow
返回

Coresets for kernel clustering

delete2024-04-22
delete0
PRE
AI
S
Shaofeng H.-C. Jiang *
R
Robert Krauthgamer
J
Jianing Lou
Y
Yubo Zhang
DOI:10.1007/s10994-024-06540-zdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
We devise coresets for kernel k\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$k$$\end{document}-Means with a general kernel, and use them to obtain new, more efficient, algorithms. Kernel k\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$k$$\end{document}-Means has superior clustering capability compared to classical k\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$k$$\end{document}-Means, particularly when clusters are non-linearly separable, but it also introduces significant computational challenges. We address this computational issue by constructing a coreset, which is a reduced dataset that accurately preserves the clustering costs. Our main result is a coreset for kernel k\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$k$$\end{document}-Means that works for a general kernel and has size poly(k & varepsilon;-1)\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${{\,\textrm{poly}\,}}(k\epsilon <^>{-1})$$\end{document}. Our new coreset both generalizes and greatly improves all previous results; moreover, it can be constructed in time near-linear in n. This result immediately implies new algorithms for kernel k\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$k$$\end{document}-Means, such as a (1+& varepsilon;)\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$(1+\epsilon )$$\end{document}-approximation in time near-linear in n, and a streaming algorithm using space and update time poly(k & varepsilon;-1logn)\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${{\,\textrm{poly}\,}}(k \epsilon <^>{-1} \log n)$$\end{document}. We validate our coreset on various datasets with different kernels. Our coreset performs consistently well, achieving small errors while using very few points. We show that our coresets can speed up kernel K-MEANS++\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\textsc {k-Means++}$$\end{document} (the kernelized version of the widely used K-MEANS++\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\textsc {k-Means++}$$\end{document} algorithm), and we further use this faster kernel K-MEANS++\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\textsc {k-Means++}$$\end{document} for spectral clustering. In both applications, we achieve significant speedup and a better asymptotic growth while the error is comparable to baselines that do not use coresets.
Keyword:
Clustering
Kernel method
Coreset
PTAS

期刊

Machine Learning 封面图
Machine Learning
IF:
2.9
论文数:
2.7K
被引数:
3.4W

机构

W
Weizmann Institute of Science
学者数:
1.3W
论文数: 1.1W
被引数: 2.3W
P
peking university
学者数:
11.9W
论文数: 8.7W
被引数: 146
引用论文

引用论文

err分享
err收藏
err分享
err收藏
Nutritional Support in Inflammatory Bowel Disease
err2008-11-04
err0
PREAI
errFranz Zurita; Donald E. Rawls; Walter P. Dyck
err分享
err收藏
Can primary care research be conducted more efficiently using routinely reported practice-level data: a cluster randomised controlled trial conducted in England?
err2022-07-01
err0
errOAAI
errPeter S Blair; Jenny Ingram; Clare Clement; Grace Young; Penny Seume; Jodi Taylor; Christie Cabral; Patricia Jane Lucas; Elizabeth Beech; Jeremy Horwood; Padraig Dixon; Martin C Gulliford; Nick Francis; Sam T Creavin; Athene Lane; Scott Bevan; Alastair D Hay
err分享
err收藏
学者 查看更多内容