返回
KD-Crowd: a knowledge distillation framework for learning from crowds
DOI:10.1007/s11704-023-3578-7.png)
摘要
En 中文
Recently, crowdsourcing has established itself as an efficient labeling solution by distributing tasks to crowd workers. As the workers can make mistakes with diverse expertise, one core learning task is to estimate each worker's expertise, and aggregate over them to infer the latent true labels. In this paper, we show that as one of the major research directions, the noise transition matrix based worker expertise modeling methods commonly overfit the annotation noise, either due to the oversimplified noise assumption or inaccurate estimation. To solve this problem, we propose a knowledge distillation framework (KD-Crowd) by combining the complementary strength of noise-model-free robust learning techniques and transition matrix based worker expertise modeling. The framework consists of two stages: in Stage 1, a noise-model-free robust student model is trained by treating the prediction of a transition matrix based crowdsourcing teacher model as noisy labels, aiming at correcting the teacher's mistakes and obtaining better true label predictions; in Stage 2, we switch their roles, retraining a better crowdsourcing model using the crowds' annotations supervised by the refined true label predictions given by Stage 1. Additionally, we propose one f-mutual information gain (MIGf) based knowledge distillation loss, which finds the maximum information intersection between the student's and teacher's prediction. We show in experiments that MIGf achieves obvious improvements compared to the regular KL divergence knowledge distillation loss, which tends to force the student to memorize all information of the teacher's prediction, including its errors. We conduct extensive experiments showing that, as a universal framework, KD-Crowd substantially improves previous crowdsourcing methods on true label prediction and worker expertise estimation.
Keyword:
crowdsourcing
label noise
worker expertise
knowledge distillation
robust learning
期刊
IF:
4.6
论文数:
1.6K
被引数:
2.8K
机构
暂无机构信息
引用论文
Exploiting Cross-Modal Prediction and Relation Consistency for Semisupervised Image Captioning利用跨模态预测和关系一致性进行半监督图像字幕
Wallerian degeneration after spinal cord lesions in cats detected with diffusion tensor imaging
NeuroImage
IF0
To Be Seen or to Hide: Visual Characteristics of Body Patterns for Camouflage and Communication in the Australian Giant CuttlefishSepia apama被看到或隐藏: 澳大利亚巨型墨鱼 Sepia apama 中伪装和交流的身体图案的视觉特征
没有更多内容

