返回
Optimizing Text Quantifiers for Multivariate Loss Functions
DOI:10.1145/2700406.png)
摘要
En 中文
We address the problem of quantification, a supervised learning task whose goal is, given a class, to estimate the relative frequency (or prevalence) of the class in a dataset of unlabeled items. Quantification has several applications in data and text mining, such as estimating the prevalence of positive reviews in a set of reviews of a given product or estimating the prevalence of a given support issue in a dataset of transcripts of phone calls to tech support. So far, quantification has been addressed by learning a general-purpose classifier, counting the unlabeled items that have been assigned the class, and tuning the obtained counts according to some heuristics. In this article, we depart from the tradition of using general-purpose classifiers and use instead a supervised learning model for structured prediction, capable of generating classifiers directly optimized for the (multivariate and nonlinear) function used for evaluating quantification accuracy. The experiments that we have run on 5,500 binary high-dimensional datasets (averaging more than 14,000 documents each) show that this method is more accurate, more stable, and more efficient than existing state-of-the-art quantification methods.
Keyword:
Algorithm
Design
Experimentation
Measurements
Quantification
prevalence estimation
prior estimation
supervised learning
text classification
loss functions
Kullback-Leibler divergence
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
4.8
论文数:
1.3K
被引数:
4.4K
机构
引用论文
Partially identified prevalence estimation under misclassification using the kappa coefficient使用kappa系数在错误分类下部分识别的患病率估计
A response to Webb and Ting's On the application of ROC analysis to predict classification performance under varying class distributions
MACHINE LEARNING
IF2.9

