arrow
Return

Promoting Unsupervised Data-To-Text Generation Using Retraining and Unified Linearization

delete2025-10-25
delete0
PRE
AI
X
Xiaobo Wang
X
Xuan Zhang *
J
Jing Cheng
K
Kunpeng Du
高宸 (Chen Gao)
M
Ma, Zhuxian
L
Liu, Bo
DOI:10.1002/cpe.70254delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In recent years, many studies have focused on unsupervised data-to-text generation methods. However, existing unsupervised methods still require a large amount of unlabeled sample training, leading to significant data collection overhead. We propose a low-resource unsupervised method called CycleRUR. This method first converts various forms of structured data (such as tables, knowledge graph(KG) triples, and meaning representations(MR)) into unified KG triples to improve the model's ability to adapt to different structured data. Additionally, CycleRUR incorporates a retraining module and a contrastive learning module within a cycle training framework, enabling the model to learn and converge from a small amount of unpaired KG triples and reference text corpus, thereby improving the model's accuracy and convergence speed. We evaluated the model's performance on the WebNLG and E2E datasets. Using only 10% of unpaired training data, our method achieved the effects of fully supervised fine-tuning. On the WebNLG dataset, it resulted in an 18.41% improvement in METEOR compared to supervised models. On the E2E dataset, it achieved improvements of 1.37% in METEOR and 4.97% in BLEU. Experiments also demonstrated that under unified linearization, CycleRUR exhibits good generalization capabilities.
Keywords:
contrastive learning
cycle training
data-to-text
retraining
unified linearization
unsupervised

Journal

C
CONCURRENCY AND COMPUTATION-PRACTICE & EXPERIENCE
IF:
1.5
Papers:
473
Citations:
0

Organization

Y
Yunnan University
Scholars:
1.6W
Papers: 9.9K
Citations: 13