返回
Improved Machine Reading Comprehension Using Data Validation for Weakly Labeled Data
DOI:10.1109/ACCESS.2019.2963569.png)
摘要
En 中文
Machine reading comprehension (MRC) is a natural language processing task wherein a given question is answered according to a holistic understanding of a given context. Recently, many researchers have shown interest in MRC, for which a considerable number of datasets are being released. Datasets for MRC, which are composed of the context-query-answer triple, are designed to answer a given query by referencing and understanding a readily-available, relevant context text. The TriviaQA dataset is a weakly labeled dataset, because it contains irrelevant context that forms no basis for answering the query. The existing syntactic data cleaning method struggles to deal with the contextual noise this irrelevancy creates. Therefore, a semantic data cleaning method using reasoning processes is necessary. To address this, we propose a new MRC model in which the TriviaQA dataset is validated and trained using a high-quality dataset. The data validation method in our MRC model improves the quality of the training dataset, and the answer extraction model learns with the validated training data, because of our validation method. Our proposed method showed a 4.33% improvement in performance for the TriviaQA Wiki, compared to the existing baseline model. Accordingly, our proposed method can address the limitation of irrelevant context in MRC better than the human supervision.
Keyword:
Computational and artificial intelligence
data validation
natural language processing
neural networks
machine reading comprehension
weak label
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Chloroplast pH values and buffer capacities in darkened leaves as revealed by CO2 solubilization in vivo
Planta
IF0
ComQA: Question Answering Over Knowledge Base via Semantic MatchingComQA: 通过语义匹配对知识库进行问答
IEEE ACCESS
IF3.6
Machine Learning Based Optimized Pruning Approach for Decoding in Statistical Machine Translation统计机器翻译中基于机器学习的优化剪枝解码方法
IEEE ACCESS
IF3.6

