Return
Feature selection based on term frequency deviation rate for text classification
DOI:10.1007/s10489-020-01937-4.png)
Abstract
En 中文
Feature selection is a technique to select a subset of the most relevant features for modeling training. In this paper, a new concept of TDR is firstly proposed to improve the classification accuracy. Then, a TDR-based algorithm for text classification is advanced. Finally, the extensive experiments are made on seven datasets (K1a, K1b, WAP, R52, R8, 20NewGroups, and Cade12) for two classifiers of Naive Bayes and Support Vector Machine. The experimental results indicate that the new approach can improve the classification accuracy by an average percent of 7.9%.
Keywords:
Text classification
Feature selection
Term frequency
Document frequency
Deviation ratio
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
3.5
Papers:
7.5K
Citations:
1.7W
Organization
No organization information available

