arrow
返回

Improving machine translation performance by exploiting non-parallel corpora

delete2005-12-01
delete179
delete
OA
AI
M
Munteanu, DS
M
Marcu, D
DOI:10.1162/089120105775299168delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
We present a novel method for discovering parallel sentences in comparable, non-parallel corpora. We train a maximum entropy classifier that, given a pair of sentences, can reliably determine whether or not they are translations of each other. Using this approach, we extract parallel data from large Chinese, Arabic, and English non-parallel newspaper corpora. We evaluate the quality of the extracted data by showing that it improves the performance of a state-of-the-art statistical machine translation system. We also show that a good-quality MT system can be built from scratch by starting with a very small parallel corpus (100,000 words) and exploiting a large non-parallel corpus. Thus, our method can be applied with great benefit to language pairs for which only scarce resources are available.
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Computational Linguistics 封面图
Computational Linguistics
IF:
5.3
论文数:
837
被引数:
2.7K

机构

暂无机构信息
引用论文

引用论文

Direct measurements of two-way wave-particle energy transfer in a collisionless space plasma
err2018-09-07
err0
errOAAI
errN. Kitamura; M. Kitahara; M. Shoji; Y. Miyoshi; H. Hasegawa; S. Nakamura; Y. Katoh; Y. Saito; S. Yokota; D. J. Gershman; A. F. Vinas; B. L. Giles; T. E. Moore; W. R. Paterson; C. J. Pollock; C. T. Russell; R. J. Strangeway; S. A. Fuselier; J. L. Burch
err分享
err收藏
The Web as a parallel corpus
err2003-09-01
err284
errOAAI
errResnik, P; Smith, NA
err分享
err收藏