Return
DOREMI: Optimizing Long Tail Predictions in Document-Level Relation Extraction
DOI:10.1016/j.knosys.2026.115359.png)
Abstract
En 中文
• We propose DOcument-level Relation Extraction optiMizing the long taIl (DOREMI), an iterative system tailored for long-tail relations to enhance the distantly supervised dataset through disagreement-driven annotations which, to our knowledge, is the first DocRE model featuring a human-in-the-loop strategy for denoising. • We demonstrate that measuring the disagreement between multiple models is a good proxy to identify Hard-To-Classify examples and yields substantial performance improvements with negligible human effort. • We release two Denoised Distantly Supervised Datasets (DDSs), one based on DocRED and one on Re-DocRED, which can be used to train any DocRE model. These DDSs greatly improve the prediction of long-tail relations and complement existing denoising approaches, as confirmed by our experimental evaluation.
Keywords:
Document-Level Relation Extraction
Active Learning
Long-Tail Relations
Natural Language Processing
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W

