arrow
返回

An enhanced Support Vector Machine classification framework by using Euclidean distance function for text document categorization

delete2011-08-25
delete108
PRE
AI
L
Lam Hong Lee *
D
Dino Isa
DOI:10.1007/s10489-011-0314-zdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This paper presents the implementation of a new text document classification framework that uses the Support Vector Machine (SVM) approach in the training phase and the Euclidean distance function in the classification phase, coined as Euclidean-SVM. The SVM constructs a classifier by generating a decision surface, namely the optimal separating hyper-plane, to partition different categories of data points in the vector space. The concept of the optimal separating hyper-plane can be generalized for the non-linearly separable cases by introducing kernel functions to map the data points from the input space into a high dimensional feature space so that they could be separated by a linear hyper-plane. This characteristic causes the implementation of different kernel functions to have a high impact on the classification accuracy of the SVM. Other than the kernel functions, the value of soft margin parameter, C is another critical component in determining the performance of the SVM classifier. Hence, one of the critical problems of the conventional SVM classification framework is the necessity of determining the appropriate kernel function and the appropriate value of parameter C for different datasets of varying characteristics, in order to guarantee high accuracy of the classifier. In this paper, we introduce a distance measurement technique, using the Euclidean distance function to replace the optimal separating hyper-plane as the classification decision making function in the SVM. In our approach, the support vectors for each category are identified from the training data points during training phase using the SVM. In the classification phase, when a new data point is mapped into the original vector space, the average distances between the new data point and the support vectors from different categories are measured using the Euclidean distance function. The classification decision is made based on the category of support vectors which has the lowest average distance with the new data point, and this makes the classification decision irrespective of the efficacy of hyper-plane formed by applying the particular kernel function and soft margin parameter. We tested our proposed framework using several text datasets. The experimental results show that this approach makes the accuracy of the Euclidean-SVM text classifier to have a low impact on the implementation of kernel functions and soft margin parameter C.
Keyword:
Text document classification
Support Vector Machine
Euclidean distance function
Kernel function, Soft margin parameter
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

U
universiti tunku abdul rahman (utar)
学者数:
2.2K
论文数: 1.8K
被引数: 2
引用论文

引用论文

Building a qualitative recruitment system via SVM with MCDM approach
err2010-01-29
err11
PREAI
errLi, Yung-Ming; Lai, Cheng-Yang; Kao, Chien-Pang
err分享
err收藏
An intelligent system for automated breast cancer diagnosis and prognosis using SVM based classifiers
err2007-07-12
err125
PREAI
errMaglogiannis, Ilias; Zafiropoulos, Elias; Anagnostopoulos, Ioannis
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容