1
Return

Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models

delete2025-01-07
delete0
delete
OA
AI
J
Jianhui Pang
F
Fanghua Ye
D
Derek F. Wong *
D
Dian Yu
施树明 cover
施树明 (Shuming Shi)
Z
Zhaopeng Tu
L
Longyue Wang *
DOI:10.1162/tacl_a_00730delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The evolution of Neural Machine Translation (NMT) has been significantly influenced by six core challenges (Koehn and Knowles, 2017) that have acted as benchmarks for progress in this field. This study revisits these challenges, offering insights into their ongoing relevance in the context of advanced Large Language Models (LLMs): domain mismatch, amount of parallel data, rare word prediction, translation of long sentences, attention model as word alignment, and sub-optimal beam search. Our empirical findings show that LLMs effectively reduce reliance on parallel data for major languages during pretraining and significantly improve translation of long sentences containing approximately 80 words, even translating documents up to 512 words. Despite these improvements, challenges in domain mismatch and rare word prediction persist. While NMT-specific challenges like word alignment and beam search may not apply to LLMs, we identify three new challenges in LLM-based translation: inference efficiency, translation of low-resource languages during pretraining, and human-aligned evaluation.

Journal

T
Transactions of the Association for Computational Linguistics
IF:
6.9
Papers:
486
Citations:
5.7K

Organization

U
University College London
Scholars:
7.9W
Papers: 6.2W
Citations: 15.7W
U
university of london
Scholars:
21.3W
Papers: 19.6W
Citations: 302
U
University of Macau
Scholars:
1.1W
Papers: 1.3W
Citations: 2.0W
Cited Papers

Cited Papers

Citing Papers

Citing Papers