返回
Using the absolute difference of term occurrence probabilities in binary text categorization
DOI:10.1007/s10489-010-0250-3.png)
摘要
En 中文
In this study, the differences among widely used weighting schemes are studied by means of ordering terms according to their discriminative abilities using a recently developed framework which expresses term weights in terms of the ratio and absolute difference of term occurrence probabilities. Having observed that the ordering of terms is dependent on the weighting scheme under concern, it is emphasized that this can be explained by the way different schemes use term occurrence differences in generating term weights. Then, it is proposed that the relevance frequency which is shown to provide the best scores on several datasets can be improved by taking into account the way absolute difference values are used in other widely used schemes. Experimental results on two different datasets have shown that improved F-1 scores can be achieved.
Keyword:
Term occurrence probability
Term weighting
Relevance frequency
Mutual information
Chi-square
Odds ratio
Text categorization
期刊
IF:
3.5
论文数:
7.6K
被引数:
1.7W
机构
引用论文
A hierarchical neural network document classifier with linguistic feature selection
APPLIED INTELLIGENCE
IF3.5
On machine learning methods for Chinese document categorization面向中文文档分类的机器学习方法研究
APPLIED INTELLIGENCE
IF3.5
Robust classification for spam filtering by back-propagation neural networks using behavior-based features使用基于行为的特征通过反向传播神经网络进行垃圾邮件过滤的鲁棒分类
APPLIED INTELLIGENCE
IF3.5

