返回
Agreement/disagreement based crowd labeling
DOI:10.1007/s10489-014-0516-2.png)
摘要
En 中文
In many supervised learning problems, determining the true labels of training instances is expensive, laborious, and even practically impossible. As an alternative approach, it is much easier to collect multiple subjective (possibly noisy) labels from human labelers, especially with the crowdsourcing services such as Amazon's Mechanical Turk. The collected labels are then aggregated to estimate the true labels. In order to reduce the negative effects of novices, spammers, and malicious labelers, it necessitates taking into account the accuracies of the labelers. However, in the absence of true labels, we miss the main source of information to estimate the labeler accuracies. This paper demonstrates that the agreements or disagreements among labeler opinions are useful sources of information and facilitate the accuracy estimation problem. We represent this estimation problem as an optimization problem which its goal is to minimize the differences between the analytical probabilities of disagreements based on estimated accuracies and the probabilities of disagreements according to the provided labels. We present an efficient semi-exhaustive search method to solve this optimization problem. Our experiments on the simulated data and three real datasets show that the proposed method is a promising idea in this emerging new area. The source code of the proposed method is available for downloading at http://ceit.aut.ac.ir/similar to amirkhani.
Keyword:
Artificial intelligence
Supervised learning
Noisy labels
Human labelers
Crowdsourcing
Agreement/disagreement
期刊
IF:
3.5
论文数:
7.6K
被引数:
1.7W
机构
引用论文
Identifying mislabeled training data with the aid of unlabeled data借助未标记数据识别错误标记的训练数据
APPLIED INTELLIGENCE
IF3.5
Efficiently Scaling up Crowdsourced Video Annotation A Set of Best Practices for High Quality, Economical Video Labeling有效地扩展众包视频注释一组用于高质量,经济的视频标记的最佳实践
Cancer classification using ensemble of neural networks with multiple significant gene subset's基于多组显著基因子集的神经网络集成模型的癌症分类
APPLIED INTELLIGENCE
IF3.5

