返回
Vertical Ensemble Co-Training for Text Classification
DOI:10.1145/3137114.png)
摘要
En 中文
High-quality, labeled data is essential for successfully applying machine learning methods to real-world text classification problems. However, in many cases, the amount of labeled data is very small compared to that of the unlabeled, and labeling additional samples could be expensive and time consuming. Co-training algorithms, which make use of unlabeled data to improve classification, have proven to be very effective in such cases. Generally, co-training algorithms work by using two classifiers, trained on two different views of the data, to label large amounts of unlabeled data. Doing so can help minimize the human effort required for labeling new data, as well as improve classification performance. In this article, we propose an ensemble-based co-training approach that uses an ensemble of classifiers from different training iterations to improve labeling accuracy. This approach, which we call vertical ensemble, incurs almost no additional computational cost. Experiments conducted on six textual datasets show a significant improvement of over 45% in AUC compared with the original co-training algorithm.
Keyword:
Co-training
text classification
ensemble
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.6
论文数:
1.5K
被引数:
6.2K
机构
引用论文
Obstructive sleep apnea classification based on spectrogram patterns in the electrocardiogram基于心电图谱图模式的阻塞性睡眠呼吸暂停分类
Coalition game for user association and bandwidth allocation in ultra-dense mmWave networks超密集毫米波网络中用户关联和带宽分配的联盟博弈

