返回
Online multi-label dependency topic models for text classification
DOI:10.1007/s10994-017-5689-6.png)
摘要
En 中文
Multi-label text classification is an increasingly important field as large amounts of text data are available and extracting relevant information is important in many application contexts. Probabilistic generative models are the basis of a number of popular text mining methods such as Naive Bayes or Latent Dirichlet Allocation. However, Bayesian models for multi-label text classification often are overly complicated to account for label dependencies and skewed label frequencies while at the same time preventing overfitting. To solve this problem we employ the same technique that contributed to the success of deep learning in recent years: greedy layer-wise training. Applying this technique in the supervised setting prevents overfitting and leads to better classification accuracy. The intuition behind this approach is to learn the labels first and subsequently add a more abstract layer to represent dependencies among the labels. This allows using a relatively simple hierarchical topic model which can easily be adapted to the online setting. We show that our method successfully models dependencies online for large-scale multi-label datasets with many labels and improves over the baseline method not modeling dependencies. The same strategy, layer-wise greedy training, also makes the batch variant competitive with existing more complex multi-label topic models.
Keyword:
Multi-label classification
Online learning
LDA
Topic model
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
2.9
论文数:
2.7K
被引数:
3.4W
机构
引用论文
Morphology and properties of melt-spun polycarbonate fibers containing single- and multi-wall carbon nanotubes
Polymer
IF0
Hierarchical Multi-label Classification using Fully Associative Ensemble Learning使用全关联集成学习的分层多标签分类
PATTERN RECOGNITION
IF7.6

