Return
Feature selection based on absolute deviation factor for text classification
DOI:10.1016/j.ipm.2022.103251.png)
Abstract
En 中文
In text classification, it is necessary to perform feature selection to alleviate the curse of dimensionality caused by high-dimensional text data. In this paper, we utilize class term frequency (CTF) and class document frequency (CDF) to characterize the relevance between terms and categories in the level of term frequency (TF) and document frequency (DF). On the basis of relevance measurement above, three feature selection methods (ADF based on CTF (ADF-CTF), ADF based on CDF (ADF-CDF), and ADF based on both CTF and CDF (ADF-CTDF)) are proposed to identify relevant and discriminant terms by introducing absolute deviation factors (ADFs). Absolute deviation, a statistic concept, is first adopted to measure the relevance divergence characterized by CTF and CDF. In addition, ADF-CTF and ADF-CDF can be combined with existing DF-based and TF-based methods, respectively, which results in new ADF-based methods. Experimental results on six high-dimensional textual datasets using three classifiers indicate that ADF-based methods outperform original DF-based and TF-based ones in 89% cases in terms of Micro-F1 and Macro-F1, which demonstrates the role of ADF integrated in existing methods to boost the classification performance. In addition, findings also show that ADF-CTDF ranks first averagely among multiple datasets and significantly outperforms other methods in 99% cases.
Keywords:
Text classification
Feature selection
Term frequency
Document frequency
Absolute deviation factor
Journal
I
IF:
6.9
Papers:
5.2K
Citations:
1.4W
Organization
Cited Papers
RETRACTED: Automatic text classification using machine learning and optimization algorithms (Retracted Article)
SOFT COMPUTING
IF2.5
Toxoplasmosis infection among pregnant women in Africa: A systematic review and meta-analysis
PLOS ONE
IF0

