arrow
Return

Unsupervised statistical text simplification using pre-trained language modeling for initialization

delete2022-08-08
delete12
PRE
AI
强继朋 cover
强继朋 (Jipeng Qiang)
F
Feng Zhang
Y
Yun Li *
Y
Yunhao Yuan
Y
Yi Zhu
X
Xindong Wu
DOI:10.1007/s11704-022-1244-0delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Unsupervised text simplification has attracted much attention due to the scarcity of high-quality parallel text simplification corpora. Recent an unsupervised statistical text simplification based on phrase-based machine translation system (UnsupPBMT) achieved good performance, which initializes the phrase tables using the similar words obtained by word embedding modeling. Since word embedding modeling only considers the relevance between words, the phrase table in UnsupPBMT contains a lot of dissimilar words. In this paper, we propose an unsupervised statistical text simplification using pre-trained language modeling BERT for initialization. Specifically, we use BERT as a general linguistic knowledge base for predicting similar words. Experimental results show that our method outperforms the state-of-the-art unsupervised text simplification methods on three benchmarks, even outperforms some supervised baselines.
Keywords:
text simplification
pre-trained language modeling
BERT
word embeddings

Journal

Frontiers of Computer Science cover
Frontiers of Computer Science
IF:
4.6
Papers:
1.6K
Citations:
2.8K

Organization

H
hefei university of technology
Scholars:
2.5W
Papers: 1.7W
Citations: 35
Y
Yangzhou University
Scholars:
2.8W
Papers: 1.9W
Citations: 3.3W