arrow
Return

Text classification using a few labeled examples

delete2014-01-01
delete36
PRE
AI
F
Francesco Colace *
M
Massimo De Santo
L
Luca Greco
P
Paolo Napoletano
DOI:10.1016/j.chb.2013.07.043delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Supervised text classifiers need to learn from many labeled examples to achieve a high accuracy. However, in a real context, sufficient labeled examples are not always available because human labeling is enormously time-consuming. For this reason, there has been recent interest in methods that are capable of obtaining a high accuracy when the size of the training set is small. In this paper we introduce a new single label text classification method that performs better than baseline methods when the number of labeled examples is small. Differently from most of the existing methods that usually make use of a vector of features composed of weighted words, the proposed approach uses a structured vector of features, composed of weighted pairs of words. The proposed vector of features is automatically learned, given a set of documents, using a global method for term extraction based on the Latent Dirichlet Allocation implemented as the Probabilistic Topic Model. Experiments performed using a small percentage of the original training set (about 1%) confirmed our theories. (C) 2013 Elsevier Ltd. All rights reserved.
Keywords:
Text mining
Text classification
Term extraction
Probabilistic topic
Model
Data mining
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Computers in Human Behavior cover
Computers in Human Behavior
IF:
8.9
Papers:
9.0K
Citations:
5.8W

Organization

U
University of Salerno
Scholars:
1.2W
Papers: 1.1W
Citations: 1.2W
U
University of Milan
Scholars:
5.1W
Papers: 3.9W
Citations: 5.0W