arrow
Return

Bilingual recursive neural network based data selection for statistical machine translation

delete2016-09-01
delete16
PRE
AI
D
Derek F. Wong
卢艺 (Yi Lü)
L
Lidia S. Chao *
DOI:10.1016/j.knosys.2016.05.003delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Data selection is a widely used and effective solution to domain adaptation in statistical machine translation (SMT). The dominant methods are perplexity-based ones, which do not consider the mutual translations of sentence pairs and tend to select short sentences. In this paper, to address these problems, we propose bilingual semi-supervised recursive neural network data selection methods to differentiate domain-relevant data from out-domain data. The proposed methods are evaluated in the task of building domain-adapted SMT systems. We present extensive comparisons and show that the proposed methods outperform the state-of-the-art data selection approaches. (C) 2016 Elsevier B.V. All rights reserved.
Keywords:
Data selection
Machine translation
Domain adaptation
Recursive neural network
Autoencoder
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

U
University of Macau
Scholars:
1.1W
Papers: 1.3W
Citations: 2.0W