返回
Ternary encoding based feature extraction for binary text classification
DOI:10.1007/s10489-014-0515-3.png)
摘要
En 中文
A novel framework for termset based feature extraction is proposed for binary text classification. The proposed approach is based on the encoding of the terms within a termset. The ternary codes '+1' and '-1' are used to represent the class that the term supports, whereas '0' denotes no support to any of the classes. Four different encoding schemes are proposed where the term weights and the term occurrence probabilities in the positive and negative documents are used to define the ternary code of a given term. The ternary patterns are utilized to define novel features by splitting them into positive and negative codes where each code is treated as a different feature extractor. Use of the derived features individually and together with bag of words representation are both investigated. The histograms of the resultant features are also employed to study the improvements that can be achieved using a small number of additional features to augment bag of words representation. Experiments conducted on four benchmark datasets with different characteristics have shown that the proposed feature extraction framework provides significant improvements compared to the bag of words representation.
Keyword:
Local ternary patterns
Feature extraction
Termsets
n-grams
Termset weighting
Text classification
期刊
IF:
3.5
论文数:
7.6K
被引数:
1.7W
机构
引用论文
An enhanced Support Vector Machine classification framework by using Euclidean distance function for text document categorization
APPLIED INTELLIGENCE
IF3.5
Robust classification for spam filtering by back-propagation neural networks using behavior-based features使用基于行为的特征通过反向传播神经网络进行垃圾邮件过滤的鲁棒分类
APPLIED INTELLIGENCE
IF3.5

