arrow
返回

Text classification based on multi-word with support vector machine

delete2008-12-01
delete195
PRE
AI
W
Wen Zhang *
T
Taketoshi Yoshida
X
Xijin Tang
DOI:10.1016/j.knosys.2008.03.044delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
One of the main themes which support text mining is text representation: that is, its task is to look for appropriate terms to transfer documents into numerical vectors. Recently, many efforts have been invested on this topic to enrich text representation using vector space model (VSM) to improve the performances of text mining techniques such as text classification and text clustering. The main concern in this paper is to investigate the effectiveness of using multi-words for text representation on the performances of text classification. Firstly, a practical method is proposed to implement the multi-word extraction from documents based on the syntactical structure. Secondly, two strategies as general concept representation and subtopic representation are presented to represent the documents using the extracted multi-words. In particular, the dynamic k-mismatch is proposed to determine the presence of a long multi-word which is a subtopic of the content of a document. Finally, we carried out a series of experiments on classifying the Reuters-21578 documents using the representations with multi-words. We used the performance of representation in individual words as the baseline, which has the largest dimension of feature set for representation without linguistic preprocessing. Moreover, linear kernel and non-linear polynomial kernel in support vector machines (SVM) are examined comparatively for classification to investigate the effect of kernel type on their performances. Index terms with low information gain (IG) are removed from the feature set at different percentages to observe the robustness of each classification method. Our experiments demonstrate that in multi-word representation, subtopic representation outperforms the general concept representation and the linear kernel outperforms the non-linear kernel of SVM in classifying the Reuters data. The effect of applying different representation strategies is greater than the effect of applying the different SVM kernels on classification performance. Furthermore, the representation using individual words outperforms any representation using multi-words. This is consistent with the major opinions concerning the role of linguistic preprocessing on documents' features when using SVM for text classification. (C) 2008 Elsevier B.V. All rights reserved.
Keyword:
Text classification
Multi-word
Feature selection
Information gain
Support vector machine
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.2W
被引数:
4.5W

机构

J
japan advanced institute of science & technology (jaist)
学者数:
2.0K
论文数: 1.9K
被引数: 0
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
引用论文

引用论文

Technology, Intellectual Property Law and Culture
err
IF0
err2024-04-30
err0
PREAI
errMegan Rae Blakely
err分享
err收藏
Mechanism of neutral carrier mediated ion transport through ion-selective bulk membranes
err2002-05-01
err0
PREAI
errA. P. Thoma; A. Viviani-Nauer; S. Arvanitis; W. E. Morf; W. Simon
err分享
err收藏
Towards automated classification of intensive care nursing narratives
err2007-12-01
err5
PREAI
errHiissa, Marketta; Pahikkala, Tapio; Suominen, Hanna; Lehtikunnas, Tuija; Back, Barbro; Karsten, Helena; Salantera, Sanna; Salakoski, Tapio
err分享
err收藏
Isolation, culture and regeneration of avocado ( Persea americana Mill.) protoplasts
err1998-12-01
err0
PREAI
errNot Available Not Available; R. E. Litz; J. W. Grosser
err分享
err收藏
学者 查看更多内容