返回
Unsupervised statistical text simplification using pre-trained language modeling for initialization
DOI:10.1007/s11704-022-1244-0.png)
摘要
En 中文
Unsupervised text simplification has attracted much attention due to the scarcity of high-quality parallel text simplification corpora. Recent an unsupervised statistical text simplification based on phrase-based machine translation system (UnsupPBMT) achieved good performance, which initializes the phrase tables using the similar words obtained by word embedding modeling. Since word embedding modeling only considers the relevance between words, the phrase table in UnsupPBMT contains a lot of dissimilar words. In this paper, we propose an unsupervised statistical text simplification using pre-trained language modeling BERT for initialization. Specifically, we use BERT as a general linguistic knowledge base for predicting similar words. Experimental results show that our method outperforms the state-of-the-art unsupervised text simplification methods on three benchmarks, even outperforms some supervised baselines.
Keyword:
text simplification
pre-trained language modeling
BERT
word embeddings
期刊
IF:
4.6
论文数:
1.6K
被引数:
2.8K
机构
引用论文
Rift and supradetachment basins during extension: insight from the Tyrrhenian rift伸展过程中的裂谷和超剥离盆地:来自第勒尼安裂谷的启示
Cobalt and Copper Composite Oxides as Efficient Catalysts for Preferential Oxidation of CO in H2-Rich Stream钴和铜复合氧化物作为H2-Rich流中CO优先氧化的有效催化剂
Chromosome duplication and ploidy level determination in African nightshadeSolanum villosumMiller非洲茄Solanum villosumMiller的染色体加倍与倍性水平鉴定

