返回
Vicinal support vector classifier using supervised kernel-based clustering
DOI:10.1016/j.artmed.2014.01.003.png)
摘要
En 中文
Objective: Support vector machines (SVMs) have drawn considerable attention due to their high generalisation ability and superior classification performance compared to other pattern recognition algorithms. However, the assumption that the learning data is identically generated from unknown probability distributions may limit the application of SVMs for real problems. In this paper, we propose a vicinal support vector classifier (VSVC) which is shown to be able to effectively handle practical applications where the learning data may originate from different probability distributions. Methods: The proposed VSVC method utilises a set of new vicinal kernel functions which are constructed based on supervised clustering in the kernel-induced feature space. Our proposed approach comprises two steps. In the clustering step, a supervised kernel-based deterministic annealing (SKDA) clustering algorithm is employed to partition the training data into different soft vicinal areas of the feature space in order to construct the vicinal kernel functions. In the training step, the SVM technique is used to minimise the vicinal risk function under the constraints of the vicinal areas defined in the SKDA clustering step. Results: Experimental results on both artificial and real medical datasets show our proposed VSVC achieves better classification accuracy and lower computational time compared to a standard SVM. For an artificial dataset constructed from non-separated data, the classification accuracy of VSVC is between 95.5% and 96.25% (using different cluster numbers) which compares favourably to the 94.5% achieved by SVM. The VSVC training time is between 8.75 s and 17.83 s (for 2-8 clusters), considerable less than the 65.0 s required by SVM. On a real mammography dataset, the best classification accuracy of VSVC is 85.7% and thus clearly outperforms a standard SVM which obtains an accuracy of only 82.1%. A similar performance improvement is confirmed on two further real datasets, a breast cancer dataset (74.01% vs. 72.52%) and a heart dataset (84.77% vs. 83.81%), coupled with a reduction in terms of learning time (32.07 s vs. 92.08 s and 25.00 s vs. 53.31 s, respectively). Furthermore, the VSVC results in the number of support vectors being equal to the specified cluster number, and hence in a much sparser solution compared to a standard SVM. Conclusion: Incorporating a supervised clustering algorithm into the SVM technique leads to a sparse but effective solution, while making the proposed VSVC adaptive to different probability distributions of the training data. (C) 2014 Elsevier B.V. All rights reserved.
Keyword:
Support vector machines
Kernel-based data clustering
Supervised deterministic annealing
Mammographic mass classification
Biomedical data classification
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.2
论文数:
2.5K
被引数:
7.8K
机构
引用论文
Evaluation of charge transfer resistance by geometrical extrapolation of the centre of semicircular impedance diagrams通过半圆阻抗图中心的几何外推法评估电荷转移电阻
Can primary care research be conducted more efficiently using routinely reported practice-level data: a cluster randomised controlled trial conducted in England?
BMJ Open
IF0
Deterministic annealing for clustering, compression, classification, regression, and related optimization problems
PROCEEDINGS OF THE IEEE
IF25.9

