返回
Learning to Weight for Text Classification
DOI:10.1109/TKDE.2018.2883446.png)
摘要
En 中文
In information retrieval (IR) and related tasks, term weighting approaches typically consider the frequency of the term in the document and in the collection in order to compute a score reflecting the importance of the term for the document. In tasks characterized by the presence of training data (such as text classification) it seems logical that the term weighting function should take into account the distribution (as estimated from training data) of the term across the classes of interest. Although supervised term weighting approaches that use this intuition have been described before, they have failed to show consistent improvements. In this article, we analyze the possible reasons for this failure, and call consolidated assumptions into question. Following this criticism, we propose a novel supervised term weighting approach that, instead of relying on any predefined formula, learns a term weighting function optimized on the training set of interest; we dub this approach Learning to Weight (LTW). The experiments that we run on several well-known benchmarks, and using different learning methods, show that our method outperforms previous term weighting approaches in text classification.
Keyword:
Training data
Task analysis
Training
Neural networks
Feature extraction
Time-frequency analysis
Information retrieval
Term weighting
supervised term weighting
text classification
neural networks
deep learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
10.4
论文数:
6.8K
被引数:
3.2W
机构
引用论文
Population Size and Migration of Anopheles gambiae in the Bancoumana Region of Mali and Their Significance for Efficient Vector Control
PLoS ONE
IF0

