arrow
返回

KD-Crowd: a knowledge distillation framework for learning from crowds

delete2024-11-11
delete2
PRE
AI
S
Shao-Yuan Li *
Y
Yuxiang Zheng
Y
Ye Shi
黄
黄圣君 (Sheng-Jun Huang)
Songcan Chen 封面图
Songcan Chen (Songcan Chen)
DOI:10.1007/s11704-023-3578-7delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Recently, crowdsourcing has established itself as an efficient labeling solution by distributing tasks to crowd workers. As the workers can make mistakes with diverse expertise, one core learning task is to estimate each worker's expertise, and aggregate over them to infer the latent true labels. In this paper, we show that as one of the major research directions, the noise transition matrix based worker expertise modeling methods commonly overfit the annotation noise, either due to the oversimplified noise assumption or inaccurate estimation. To solve this problem, we propose a knowledge distillation framework (KD-Crowd) by combining the complementary strength of noise-model-free robust learning techniques and transition matrix based worker expertise modeling. The framework consists of two stages: in Stage 1, a noise-model-free robust student model is trained by treating the prediction of a transition matrix based crowdsourcing teacher model as noisy labels, aiming at correcting the teacher's mistakes and obtaining better true label predictions; in Stage 2, we switch their roles, retraining a better crowdsourcing model using the crowds' annotations supervised by the refined true label predictions given by Stage 1. Additionally, we propose one f-mutual information gain (MIGf) based knowledge distillation loss, which finds the maximum information intersection between the student's and teacher's prediction. We show in experiments that MIGf achieves obvious improvements compared to the regular KL divergence knowledge distillation loss, which tends to force the student to memorize all information of the teacher's prediction, including its errors. We conduct extensive experiments showing that, as a universal framework, KD-Crowd substantially improves previous crowdsourcing methods on true label prediction and worker expertise estimation.
Keyword:
crowdsourcing
label noise
worker expertise
knowledge distillation
robust learning

期刊

Frontiers of Computer Science 封面图
Frontiers of Computer Science
IF:
4.6
论文数:
1.6K
被引数:
2.8K

机构

暂无机构信息
引用论文

引用论文

Multi-Label Learning from Crowds
err2019-07-01
err37
PREAI
errLi, Shao-Yuan; Jiang, Yuan; Chawla, Nitesh V.; Zhou, Zhi-Hua
err分享
err收藏
Wallerian degeneration after spinal cord lesions in cats detected with diffusion tensor imaging
err2011-08-01
err0
PREAI
errJ. Cohen-Adad; H. Leblond; H. Delivet-Mongrain; M. Martinez; H. Benali; S. Rossignol
err分享
err收藏
Learning from crowds with sparse and imbalanced annotations
err2022-06-14
err2
errOAAI
errShi, Ye; Li, Shao-Yuan; Huang, Sheng-Jun
err分享
err收藏
Crowdsourcing aggregation with deep Bayesian learning
err2021-02-07
err34
PREAI
errLi, Shao-Yuan; Huang, Sheng-Jun; Chen, Songcan
err分享
err收藏
err分享
err收藏
没有更多内容