返回
A new feature selection algorithm based on binomial hypothesis testing for spam filtering
DOI:10.1016/j.knosys.2011.04.006.png)
摘要
En 中文
Content-based spam filtering is a binary text categorization problem. To improve the performance of the spam filtering, feature selection, as an important and indispensable means of text categorization, also plays an important role in spam filtering. We proposed a new method, named Bi-Test, which utilizes binomial hypothesis testing to estimate whether the probability of a feature belonging to the spam satisfies a given threshold or not. We have evaluated Bi-Test on six benchmark spam corpora (pu1, pu2, pu3, pua, lingspam and CSDMC2010), using two classification algorithms, Naive Bayes (NB) and Support Vector Machines (SVM), and compared it with four famous feature selection algorithms (information gain, chi(2)-statistic, improved Gini index and Poisson distribution). The experiments show that Bi-Test performs significantly better than chi(2)-statistic and Poisson distribution, and produces comparable performance with information gain and improved Gini index in terms of F1 measure when Naive Bayes classifier is used; it achieves comparable performance with the other methods when SVM classifier is used. Moreover, Bi-Test executes faster than the other four algorithms. (C) 2011 Elsevier B.V. All rights reserved.
Keyword:
Feature selection
Binomial hypothesis testing
Spam filtering
Text categorization
Binomial distribution
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
K
IF:
7.6
论文数:
1.2W
被引数:
4.5W
机构
引用论文
Information gain and divergence-based feature selection for machine learning-based text categorization基于信息增益和散度的特征选择,用于基于机器学习的文本分类
M Pathway and Areas 44 and 45 Are Involved in Stereoscopic Recognition Based on Binocular Disparity.

