arrow
Return

Evaluating Document Coherence Modeling

delete2021-07-08
delete4
delete
OA
AI
A
Aili Shen *
M
Meladel Mistica
B
Bahar Salehi
H
Hang Li
T
Timothy Baldwin
J
Jianzhong Qi
DOI:10.1162/tacl_a_00388delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
While pretrained language models (LMs) have driven impressive gains over morphosyntactic and semantic tasks, their ability to model discourse and pragmatic phenomena is less clear. As a step towards a better understanding of their discourse modeling capabilities, we propose a sentence intrusion detection task. We examine the performance of a broad range of pretrained LMs on this detection task for English. Lacking a dataset for the task, we introduce INSteD, a novel intruder sentence detection dataset, containing 170,000+ documents constructed from English Wikipedia and CNN news articles. Our experiments show that pretrained LMs perform impressively in in-domain evaluation, but experience a substantial drop in the cross-domain setting, indicating limited generalization capacity. Further results over a novel linguistic probe dataset show that there is substantial room for improvement, especially in the crossdomain setting.

Journal

T
Transactions of the Association for Computational Linguistics
IF:
6.9
Papers:
486
Citations:
5.7K

Organization

U
university of melbourne
Scholars:
5.7W
Papers: 5.4W
Citations: 69