arrow
返回

Re-structuring, Re-labeling, and Re-aligning for Syntax-Based Machine Translation

delete2010-06-01
delete11
delete
OA
AI
W
Wei Wang *
J
Jonathan May
K
Kevin Knight
D
Daniel Marcu
DOI:10.1162/coli.2010.36.2.09054delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
This article shows that the structure of bilingual material from standard parsing and alignment tools is not optimal for training syntax-based statistical machine translation (SMT) systems. We present three modifications to the MT training data to improve the accuracy of a state-of-the-art syntax MT system: re-structuring changes the syntactic structure of training parse trees to enable reuse of substructures; re-labeling alters bracket labels to enrich rule application context; and re-aligning unifies word alignment across sentences to remove bad word alignments and refine good ones. Better structures, labels, and word alignments are learned by the EM algorithm. We show that each individual technique leads to improvement as measured by BLEU, and we also show that the greatest improvement is achieved by combining them. We report an overall 1.48 BLEU improvement on the NIST08 evaluation set over a strong baseline in Chinese/English translation.
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Computational Linguistics 封面图
Computational Linguistics
IF:
5.3
论文数:
837
被引数:
2.7K

机构

暂无机构信息
引用论文

引用论文

Training tree transducers
err2008-09-01
err35
errOAAI
errGraehl, Jonathan; Knight, Kevin; May, Jonathan
err分享
err收藏
err分享
err收藏
没有更多内容