arrow
返回

Separability and scatteredness (S&S) ratio-based efficient SVM regularization parameter, kernel, and kernel parameter selection

delete2025-01-27
delete0
PRE
AI
M
Mahdi Shamsi *
S
Soosan Beheshti
DOI:10.1007/s10044-025-01411-2delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Support Vector Machine (SVM) is a robust machine learning algorithm with broad applications in classification, regression, and outlier detection. SVM requires tuning a regularization parameter (RP) which controls the model capacity and the generalization performance. Conventionally, the optimum RP is found by comparison of a range of values through the Cross-Validation (CV) procedure. In addition, for non-linearly separable data, the SVM uses kernels. In this case a set of kernels, each with a set of parameters, denoted as a grid of kernels, are considered. The optimal choice of RP and the grid of kernels is through various forms of deterministic or probabilistic grid-search. The existing methods rely heavily on exhaustive searches and provide very limited insight into the underlying data characteristics, resulting in excessive computational complexity. This work addresses this issue by proposing a statistical framework that directly relates the dataset's separability and scatteredness to the choice of optimal hyperparameters. By stochastically analyzing the behavior of the regularization parameter, the method shows that the SVM performance can be modeled as a function of the newly defined separability and scatteredness (S&S) ratio of the data. The Separability is a measure of the distance between classes, and the scatteredness is the ratio of the spread of data points. In particular, for the hinge loss cost function, an S&S ratio-based table provides the optimum RP. The data S&S ratio is a powerful value that can automatically evaluate linear or non-linear separability before using the SVM algorithm. The provided lookup S&S ratio-based table can also provide the optimum kernel and its parameters before using the SVM algorithm. Consequently, the computational complexity of the CV grid-search is reduced to only the computational complexity of one-time use of the SVM. The simulation results on the real dataset confirm the superiority of the proposed approach in the sense of efficiency and computational complexity over the grid-search methods. The method performs better or comparable to the existing state of-the-art methods with a significantly reduced computational cost.
Keyword:
Support Vector Machine
Regularization parameter
Machine learning
Kernel method
Hyperparameter selection

期刊

Pattern Analysis and Applications 封面图
Pattern Analysis and Applications
IF:
2
论文数:
1.9K
被引数:
1.9K

机构

T
Toronto Metropolitan University
学者数:
6.0K
论文数: 7.0K
被引数: 6.4K
引用论文

引用论文

Text categorization: past and present
err2020-09-30
err45
PREAI
errDhar, Ankita; Mukherjee, Himadri; Dash, Niladri Sekhar; Roy, Kaushik
err分享
err收藏
err分享
err收藏
A probabilistic framework for SVM regression and error bar estimation
err2002-01-01
err113
PREAI
errGao, JB; Gunn, SR; Harris, CJ; Brown, M
err分享
err收藏
err分享
err收藏
Kidney regeneration in fish
err2018-01-01
err0
errOAAI
errThomas Bates; Uta Naumann; Beate Hoppe; Christoph Englert
err分享
err收藏
Optimised one-class classification performance优化的单类分类性能
err2022-03-11
err1
errOAAI
errLenz, Oliver Urs; Peralta, Daniel; Cornelis, Chris
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Maximum margin and global criterion based-recursive feature selection
err2024-01-01
err2
PREAI
errDing, Xiaojian; Li, Yi; Chen, Shilin
err分享
err收藏
Model selection for primal SVM
err2011-04-22
err31
errOAAI
errMoore, Gregory; Bergeron, Charles; Bennett, Kristin P.
err分享
err收藏
学者 查看更多内容