arrow
返回

Using unsupervised clustering approach to train the Support Vector Machine for text classification

delete2016-10-01
delete41
PRE
AI
N
Niusha Shafiabady *
L
Lam Hong Lee
R
Raj Rajkumar
V
Vish Kallimani
N
Nik Ahmad Akram
D
Dino Isa
DOI:10.1016/j.neucom.2015.10.137delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The use of learning algorithms for text classification assumes the availability of a large amount of documents which have been organized and labeled correctly by human experts for use in the training phase. Unless the text documents in question have been in existence for some time, using an expert system is inevitable because manual organizing and labeling of thousands of groups of text documents can be a very labor intensive and intellectually challenging activity. Also, in some new domains, the knowledge to organize and label different classes might not be unavailable. Therefore unsupervised learning schemes for automatically clustering data in the training phase are needed. Furthermore, even when knowledge exists, variation is high when the subject under classification depends on personal opinions and is open to different interpretations. This paper describes a methodology which uses Self Organizing Maps (SOM) and alternatively does the automatic clustering by using the Correlation Coefficient (CorrCoef). Consequently the clusters are used as the labels to train the Support Vector Machine (SVM). Experiments and results are presented based on applying the methodology to some standard text datasets in order to verify the accuracy of the proposed scheme. We will also present results which are used to evaluate the effect that dimensionality reduction and changes in the clustering schemes have on the accuracy of the SVM. Results show that the proposed combination has better accuracy compared to training the learning machine using the expert knowledge. (C) 2016 Elsevier B.V. All rights reserved.
Keyword:
Unsupervised learning
Classification
Support Vector Machines
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

Q
quest international university perak
学者数:
102
论文数: 84
被引数: 0
U
Universiti Teknologi Petronas
学者数:
5.4K
论文数: 4.6K
被引数: 5.9K
引用论文

引用论文

err分享
err收藏
Preparation of water-dispersible TiO2 nanoparticles from titanium tetrachloride using urea hydrogen peroxide as an oxygen donor
err2013-01-01
err0
errOAAI
errNaoko Watanabe; Taichi Kaneko; Yuko Uchimaru; Sayaka Yanagida; Atsuo Yasumori; Yoshiyuki Sugahara
err分享
err收藏
Using the self organizing map for clustering of text documents
err2009-07-01
err58
PREAI
errIsa, Dino; Kallimani, V. P.; Lee, Lam Hong
err分享
err收藏
Strategic considerations in the introduction of advanced manufacturing technologies in the Cypriot industry
err1998-12-01
err0
PREAI
errAndreas Efstathiades; Savvas A Tassou; Antonios Antoniou; George Oxinos
err分享
err收藏
Music Recommendation Using Content and Context Information Mining
err2010-01-01
err121
PREAI
errSu, Ja-Hwung; Yeh, Hsin-Ho; Yu, Philip S.; Tseng, Vincent S.
err分享
err收藏
学者 查看更多内容