返回
Noise Level Estimation for Model Selection in Kernel PCA Denoising
DOI:10.1109/TNNLS.2015.2388696.png)
摘要
En 中文
One of the main challenges in unsupervised learning is to find suitable values for the model parameters. In kernel principal component analysis (kPCA), for example, these are the number of components, the kernel, and its parameters. This paper presents a model selection criterion based on distance distributions (MDDs). This criterion can be used to find the number of components and the sigma(2) parameter of radial basis function kernels by means of spectral comparison between information and noise. The noise content is estimated from the statistical moments of the distribution of distances in the original dataset. This allows for a type of randomization of the dataset, without actually having to permute the data points or generate artificial datasets. After comparing the eigenvalues computed from the estimated noise with the ones from the input dataset, information is retained and maximized by a set of model parameters. In addition to the model selection criterion, this paper proposes a modification to the fixed-size method and uses the incomplete Cholesky factorization, both of which are used to solve kPCA in large-scale applications. These two approaches, together with the model selection MDD, were tested in toy examples and real life applications, and it is shown that they outperform other known algorithms.
Keyword:
Kernel principal component analysis (kPCA)
least squares support vector machines (LS-SVMs)
noise level estimation
unsupervised learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
8.9
论文数:
7.6K
被引数:
7.2W
机构
引用论文
Sparse kernel spectral clustering models for large-scale data analysis面向大规模数据分析的稀疏核谱聚类模型
NEUROCOMPUTING
IF6.5
PECVD-grown carbon nanotubes on silicon substrates with a nickel-seeded tip-growth structure具有镍种子尖端生长结构的硅衬底上的PECVD生长的碳纳米管

