Challenge: Recent advances in neural language modeling and multilingual training have prompted widespread adoption of machine translation (MT) technologies across an unprecedented range of world languages.
Approach: They propose to use a dataset to assess the impact of two state-of-the-art NMT systems, Google Translate and the multilingual mBART-50 model, on translation productivity.
Outcome: The proposed model is faster than translation from scratch, but the magnitude of productivity gains varies widely across systems and languages.

Similar Papers

Leveraging GPT-4 for Automatic Translation Post-Editing (2023.findings-emnlp)

Copied to clipboard

Challenge: Neural Machine Translation models still require translation post-editing to rectify errors and enhance quality under critical settings.
Approach: They use GPT-4 to automatically post-edit NMT outputs across several language pairs . they show that GPT4 is adept at translation post- editing, producing meaningful edits .
Outcome: The proposed translation post-editor improves on state-of-the-art language models on English-Chinese, English-German, Chinese-English and German-English language pairs.
English-Basque Statistical and Neural Machine Translation (L18-1)

Copied to clipboard

Challenge: Neural machine translation (NMT) requires large training corpora, which is problematic for low-resource languages.
Approach: They propose to use an open-domain and an IT-domain corpora to train machine translations in English-Basque.
Outcome: The proposed systems outperform OpenNMT, Moses SMT and Google Translate in English-Basque translation.
Enhancing Large Language Models for Document-Level Translation Post-Editing Using Monolingual Data (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have excellent performance in many tasks, but they still face challenges in document translation.
Approach: They propose a method that leverages the capabilities of Large Language Models to optimize document translation using only monolingual data.
Outcome: The proposed method improves translation quality and improves contextual consistency in document translation using only monolingual data.
Revisiting Low-Resource Neural Machine Translation: A Case Study (P19-1)

Copied to clipboard

Challenge: Recent research has shown that neural machine translation models are highly data-inefficient and underperform phrase-based statistical machine translation (PBSMT) in low-resource settings.
Approach: They propose to use auxiliary data to train low-resource neural machine translation systems without auxiliary monolingual or multilingual data.
Outcome: The proposed methods outperform PBSMT and other statistical machine translation models in Korean–English with minimal data.
Neural Machine Translation Quality and Post-Editing Performance (2021.emnlp-main)

Copied to clipboard

Challenge: a recent study has shown that MT post-editing can reduce translation quality and speed . a large-scale study involving 30 professional translators examined the relationship between MT performance and post-edited outputs.
Approach: They examine the relationship between MT performance and post-editing time and quality . they use neural MT of high quality to improve translation quality based on phrase-based MT .
Outcome: The proposed model is not stable predictor of time or quality, the authors say . they find that better MT systems lead to fewer changes in the sentences .
Multilingual Neural Machine Translation (2020.coling-tutorials)

Copied to clipboard

Challenge: In this tutorial, we will cover the latest advances in NMT to enhance low-resource translation.
Approach: They will cover the latest advances in NMT approaches that leverage multilingualism . they will focus on topics such as language divergence, transfer learning and pivoting .
Outcome: This tutorial will cover the latest advances in NMT to enhance low-resource translation models.
Towards Personalised and Document-level Machine Translation of Dialogue (2021.eacl-srw)

Copied to clipboard

Challenge: State-of-the-art (SOTA) neural machine translation systems translate texts at sentence level, ignoring context.
Approach: They propose to integrate extra-textual information into the translation process for the domain of dialogue extracted from TV subtitles in five languages: English, Brazilian Portuguese, German, French and Polish.
Outcome: The proposed systems translate texts at sentence level, ignoring context . there are no readily available robust evaluation metrics for them .
Tagged Back-translation Revisited: Why Does It Really Work? (2020.acl-main)

Copied to clipboard

Challenge: In this paper, we show that neural machine translation systems trained on large back-translated data overfit some of the characteristics of machine-transcribed texts.
Approach: They propose to add a tag to back-translations to help distinguish back-translated data from original parallel training data.
Outcome: The proposed tag helps the system distinguish back-translated data from original parallel training data and is as effective as a tag in high-resource training.
LangMark: A Multilingual Dataset for Automatic Post-Editing (2025.acl-long)

Copied to clipboard

Challenge: Automated post-editing (APE) aims to correct errors in machine-translated text . lack of large-scale multilingual datasets specifically tailored to NMT outputs hinders APE development .
Approach: They propose to use a human-annotated multilingual APE dataset for English translation to seven languages to address this gap.
Outcome: The proposed dataset offers both linguistic diversity and scale.
CODET: A Benchmark for Contrastive Dialectal Evaluation of Machine Translation (2024.findings-eacl)

Copied to clipboard

Challenge: Neural machine translation systems exhibit limited robustness in handling source-side linguistic variations.
Approach: They propose a dialectal benchmark to quantify the robustness of MT systems to handle source-side linguistic variations.
Outcome: The proposed benchmark demonstrates that large MT models face challenges translating dialectal variants.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations