arrow
返回

Extracting parallel fragments from comparable documents using a generative model

delete2019-01-01
delete1
PRE
AI
S
Somayeh Bakhshaei
R
Reza Safabakhsh *
S
Shahram Khadivi
DOI:10.1016/j.csl.2018.07.002delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Although parallel corpora are essential language resources for many natural language processing tasks, they are rare or even not available for many language pairs. Instead, comparable corpora are widely available and contain parallel fragments of information that can be used in applications like statistical machine translation systems. In this research, we propose a generative latent Dirichlet allocation based model for extracting parallel fragments from comparable documents without using any initial parallel data or bilingual lexicon. The experimental results show significant improvement if the extracted fragments generated by the proposed method are used for augmenting an existing parallel corpus in an statistical machine translation system. According to the human judgment, the accuracy of the proposed method for an English-Persian task is about 59.7%. Also, the out of vocabulary error rate for the same task is reduced by 28%. (C) 2018 Elsevier Ltd. All rights reserved.
Keyword:
Fragment extraction
Comparable corpora
Generative model
Statistical machine translation
Persian
English
German
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

C
Computer Speech and Language
IF:
3.4
论文数:
1.5K
被引数:
2.6K

机构

A
Amirkabir University of Technology
学者数:
1.1W
论文数: 1.1W
被引数: 1.0W
E
ebay inc.
学者数:
57
论文数: 43
被引数: 0