arrow
Return

Vertical Ensemble Co-Training for Text Classification

delete2017-10-25
delete13
PRE
AI
G
Gilad Katz *
C
Cornelia Caragea
A
Asaf Shabtai
DOI:10.1145/3137114delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
High-quality, labeled data is essential for successfully applying machine learning methods to real-world text classification problems. However, in many cases, the amount of labeled data is very small compared to that of the unlabeled, and labeling additional samples could be expensive and time consuming. Co-training algorithms, which make use of unlabeled data to improve classification, have proven to be very effective in such cases. Generally, co-training algorithms work by using two classifiers, trained on two different views of the data, to label large amounts of unlabeled data. Doing so can help minimize the human effort required for labeling new data, as well as improve classification performance. In this article, we propose an ensemble-based co-training approach that uses an ensemble of classifiers from different training iterations to improve labeling accuracy. This approach, which we call vertical ensemble, incurs almost no additional computational cost. Experiments conducted on six textual datasets show a significant improvement of over 45% in AUC compared with the original co-training algorithm.
Keywords:
Co-training
text classification
ensemble
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

ACM Transactions on Intelligent Systems and Technology cover
ACM Transactions on Intelligent Systems and Technology
IF:
6.6
Papers:
1.5K
Citations:
6.2K

Organization

B
ben gurion university
Scholars:
1.3W
Papers: 1.0W
Citations: 5
U
University of North Texas System
Scholars:
8.0K
Papers: 7.7K
Citations: 178