arrow
返回

Effective Bayesian-network-based missing value imputation enhanced by crowdsourcing

delete2020-02-01
delete19
PRE
AI
C
Chen Ye
王
王宏志 (Hongzhi Wang) *
W
Wenbo Lu
李
李建忠 (Jianzhong Li)
DOI:10.1016/j.knosys.2019.105199delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
During the process of data collection, incompleteness is one of the most serious data quality problems to deal with. Traditional imputation methods mostly rely on statistics and machine learning techniques. However, both types of methods are limited in their accuracy due to lacking enough information about the missing data. To obtain more information, recent methods resort to external sources such as knowledge bases or the worldwide web. Unfortunately, such methods may still be less helpful, since there may exist little information about the missing values in the knowledge bases, or too much noise on the web. To tackle these issues, this paper adopts crowdsourcing as the external source, where hundreds of thousands of ordinary workers on the platform can provide high-quality information based on contextual knowledge and human cognitive ability. To reduce the cost, a joint model is proposed for imputation, which integrates crowdsourcing into the process of Bayesian inference. We first construct a Bayesian network for the attributes in the dataset, then the missing attribute values are inferred by Bayesian inference. To improve the accuracy of the Bayesian inference, we outsource a small number of informative tasks to the crowd workers, where the informative tasks are selected based on uncertainty and influence. The proposed approach is evaluated with extensive experiments using real-world datasets with a simulated crowd and two real crowdsourcing platforms. The experimental results show that our approach achieves a better performance compared to other imputation approaches. (C) 2019 Elsevier B.V. All rights reserved.
Keyword:
Missing values
Bayesian network
Crowdsourcing
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.2W
被引数:
4.5W

机构

H
harbin institute of technology
学者数:
8.0W
论文数: 6.6W
被引数: 66
引用论文

引用论文

err分享
err收藏
Adjusted weight voting algorithm for random forests in handling missing values
err2017-09-01
err80
PREAI
errXia, Jing; Zhang, Shengyu; Cai, Guolong; Li, Li; Pan, Qing; Yan, Jing; Ning, Gangmin
err分享
err收藏
Crowdsourced Data Management: A Survey
err2016-09-01
err205
PREAI
errLi, Guoliang; Wang, Jiannan; Zheng, Yudian; Franklin, Michael J.
err分享
err收藏
Missing covariate data in medical research: To impute is better than to ignore
err2010-07-01
err460
PREAI
errJanssen, Kristel J. M.; Donders, A. Rogier T.; Harrell, Frank E., Jr.; Vergouwe, Yvonne; Chen, Qingxia; Grobbee, Diederick E.; Moons, Karel G. M.
err分享
err收藏
Current management of severe pelvic and perineal trauma
err2012-08-01
err0
PREAI
errC. Arvieux; F. Thony; C. Broux; F.-X. Ageron; E. Rancurel; J. Abba; J.-L. Faucheron; J.-J. Rambeaud; J. Tonetti
err分享
err收藏
The spectrum of somatic and germline NF1 mutations in NF1 patients with spinal neurofibromas
err2009-02-17
err0
PREAI
errMeena Upadhyaya; Gill Spurlock; Lan Kluwe; Nadia Chuzhanova; Emma Bennett; Nick Thomas; Abhijit Guha; Victor Mautner
err分享
err收藏
学者 查看更多内容