arrow
Return

Data augmentation for low-resource languages NMT guided by constrained sampling

delete2021-08-25
delete10
delete
OA
AI
M
Mieradilijiang Maimaiti
Y
Yang Liu *
H
Huanbo Luan
M
Maosong Sun
DOI:10.1002/int.22616delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Data augmentation (DA) is a ubiquitous approach for several text generation tasks. Intuitively, in the machine translation paradigm, especially in low-resource languages scenario, many DA methods have appeared. The most commonly used methods are building pseudocorpus by randomly sampling, omitting, or replacing some words in the text. However, previous approaches hardly guarantee the quality of augmented data. In this study, we try to augment the corpus by introducing a constrained sampling method. Additionally, we also build the evaluation framework to select higher quality data after augmentation. Namely, we use the discriminator submodel to mitigate syntactic and semantic errors to some extent. Experimental results show that our augmentation method consistently outperforms all the previous state-of-the-art methods on both small and large-scale corpora in eight language pairs from four corpora by 2.38-4.18 bilingual evaluation understudy points.
Keywords:
artificial intelligence
constrained sampling
data augmentation
low-resource languages
natural language processing
neural machine translation
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

International Journal of Intelligent Systems cover
International Journal of Intelligent Systems
IF:
3.7
Papers:
3.0K
Citations:
8.1K

Organization

T
tsinghua university
Scholars:
11.8W
Papers: 10.0W
Citations: 137