返回
Text document preprocessing with the Bayes formula for classification using the Support Vector Machine
DOI:10.1109/TKDE.2008.76.png)
摘要
En 中文
This work implements an enhanced hybrid classification method through the utilization of the naive Bayes approach and the Support Vector Machine (SVM). In this project, the Bayes formula was used to vectorize (as opposed to classify) a document according to a probability distribution reflecting the probable categories that the document may belong to. The Bayes formula gives a range of probabilities to which the document can be assigned according to a predetermined set of topics (categories) such as those found in the 20 Newsgroups data set for instance. Using this probability distribution as the vectors to represent the document, the SVM can then be used to classify the documents on a multidimensional level. The effects of an inadvertent dimensionality reduction caused by classifying using only the highest probability using the naive Bayes classifier can be overcome using the SVM by employing all the probability values associated with every category for each document. This method can be used for any data set and shows a significant reduction in training time as compared to the Lsquare method and significant improvement in the classification accuracy when compared to pure naive Bayes systems and also the TF-IDF/SVM hybrids.
Keyword:
document classification
Bayes formula
support vector machines
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
10.4
论文数:
6.8K
被引数:
3.2W
机构
暂无机构信息
引用论文
A new text categorization technique using distributional clustering and learning logic一种新的基于分布式聚类和学习逻辑的文本分类技术
Physiology and tonotopic organization of auditory receptors in the cricketGryllus bimaculatus DeGeer
没有更多内容

