arrow
返回

What Does My QA Model Know? Devising Controlled Probes Using Expert Knowledge

delete2020-12-01
delete17
delete
OA
AI
K
Kyle Richardson *
A
Ashish Sabharwal
DOI:10.1162/tacl_a_00331delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Open-domain question answering (QA) involves many knowledge and reasoning challenges, but are successful QA models actually learning such knowledge when trained on benchmark QA tasks? We investigate this via several new diagnostic tasks probing whether multiplechoice QA models know definitions and taxonomic reasoning-two skills widespread in existing benchmarks and fundamental to more complex reasoning. We introduce a methodology for automatically building probe datasets from expert knowledge sources, allowing for systematic control and a comprehensive evaluation. We include ways to carefully control for artifacts that may arise during this process. Our evaluation confirms that transformer-based multiple-choice QA models are already predisposed to recognize certain types of structural linguistic knowledge. However, it also reveals a more nuanced picture: their performance notably degrades even with a slight increase in the number of hops'' in the underlying taxonomic hierarchy, and with more challenging distractor candidates. Further, existing models are far from perfect when assessed at the level of clusters of semantically connected probes, such as all hypernym questions about a single concept.
Keyword:
SCIENCE

期刊

T
Transactions of the Association for Computational Linguistics
IF:
6.9
论文数:
486
被引数:
5.7K

机构

暂无机构信息
引用论文

引用论文

Functional Neuroimaging of Fatigue
err2009-05-01
err0
PREAI
errJohn DeLuca; Helen M. Genova; Emlyn J. Capili; Glenn R. Wylie
err分享
err收藏
Gradients of connectivity distance in the cerebral cortex of the macaque monkey
err2018-12-13
err0
errOAAI
errSabine Oligschläger; Ting Xu; Blazej M. Baczkowski; Marcel Falkiewicz; Arnaud Falchier; Gary Linn; Daniel S. Margulies
err分享
err收藏