返回
RESHAPE: Reverse-Edited Synthetic Hypotheses for Automatic Post-Editing
DOI:10.1109/ACCESS.2022.3154768.png)
摘要
En 中文
Synthetic training data has been extensively used to train Automatic Post-Editing (APE) models in many recent studies because the quantity of human-created data has been considered insufficient. However, the most widely used synthetic APE dataset, eSCAPE, overlooks respecting the minimal editing property of genuine data, and this defect may have been a limiting factor for the performance of APE models. This article suggests adapting back-translation to APE to constrain edit distance, while using stochastic sampling in decoding to maintain the diversity of outputs, to create a new synthetic APE dataset, RESHAPE. Our experiments show that (1) RESHAPE contains more samples resembling genuine APE data than eSCAPE does, and (2) using RESHAPE as new training data improves APE models' performance substantially over using eSCAPE.
Keyword:
Decoding
Data models
Training data
Training
Limiting
Feeds
Transformers
Automatic post-editing
back-translation
decoding strategy
machine translation
synthetic data generation
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
暂无机构信息
引用论文
[P3–489]: IN SUPPORT OF A NATIONAL DEMENTIA PLAN: UNDERSTANDING DEMENTIA CARE IN FILIPINO HOMES[P3-489]: 支持国家痴呆症计划: 了解菲律宾家庭中的痴呆症护理
没有更多内容

