Challenge: Recent machine translation methods are highly sensitive to orthographical variations such as spelling errors.
Approach: They propose to train machine translation models with random synthetic noise at training time . they focus on translation performance on natural typos, and show robustness to such noise .
Outcome: The proposed method significantly improves translation models on natural typos without accessing natural noise data or distribution.

Similar Papers

Improving Robustness of Machine Translation with Synthetic Noise (N19-1)

Copied to clipboard

Challenge: Recent work on MT robustness has demonstrated the need to build or adapt systems that are resilient to such noise.
Approach: They propose to synthesize natural noise in social media data to enhance robustness of MT systems by leveraging natural noise.
Outcome: The proposed method can make a vanilla MT system more resilient to noise, partially mitigating loss in accuracy resulting therefrom.
Visual Cues and Error Correction for Translation Robustness (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing robustness techniques fail when faced with unseen types of noise and their performance degrades on clean texts.
Approach: They propose visual context to improve translation robustness for noisy texts . they also propose an error correction training regime that can be used as an auxiliary task .
Outcome: The proposed training regime improves translation robustness on noisy texts while maintaining translation quality on clean texts.
Did Translation Models Get More Robust Without Anyone Even Noticing? (2025.acl-long)

Copied to clipboard

Challenge: Neural machine translation models are highly sensitive to “noisy” inputs, such as spelling errors, abbreviations, and formatting issues.
Approach: They revisit this insight in light of recent multilingual MT models and large language models applied to machine translation.
Outcome: The proposed models perform better on clean data than previous models, but none of the open models use robustness techniques.
Machine Translation Robustness to Natural Asemantic Variation (2022.emnlp-main)

Copied to clipboard

Challenge: Existing machine translation models struggle with noisy data and tail-end words and phrases.
Approach: They introduce and formalize a class of noise and variation that preserves meaning in the target language.
Outcome: The proposed model can perform better on natural asemantic variation (NAV) the proposed model is robust to a variety of perturbations, but not all of them are achieved with organic variations.
Leveraging Synthetic Targets for Machine Translation (2023.findings-acl)

Copied to clipboard

Challenge: Using synthetic target data, training models on synthetic targets outperforms training on actual ground-truth data.
Approach: They propose a recipe for training machine translation models on synthetic target data by leveraging a large pre-trained model.
Outcome: The proposed model outperforms training on real-world translation datasets.
Beyond Static Synthetic Noise: Assessing the Robustness of Large Language Models to Natural Context Variation in the Real World (2026.findings-acl)

Copied to clipboard

Challenge: Current robustness evaluation methods rely on static synthetic perturbations to stress-test models.
Approach: They propose a framework for automatically evaluating QA models under naturally occurring textual perturbations by replacing context passages with revised Wikipedia edit histories.
Outcome: The proposed framework replaces context passages with revised Wikipedia edit histories to improve model performance.
Beyond Noise: Mitigating the Impact of Fine-grained Semantic Divergences on Neural Machine Translation (2021.acl-long)

Copied to clipboard

Challenge: Prior work treats all types of mismatches between source and target as noise . Consequently, it remains unclear how noisy parallel training samples impact NMT training.
Approach: They propose a divergent-aware NMT framework that uses factors to help NMT recover from the degradation caused by naturally occurring divergences.
Outcome: The proposed framework improves translation quality and model calibration on EN-FR tasks.
Improving Neural Machine Translation Robustness via Data Augmentation: Beyond Back-Translation (D19-55)

Copied to clipboard

Challenge: Neural Machine Translation models are sensitive to noise in the input data.
Approach: They propose new methods to extend limited noisy data and further improve NMT robustness to noise while keeping the models small.
Outcome: The proposed methods extend limited noisy data and improve robustness to noise while keeping the models small.
Robust to Noise Models in Natural Language Processing Tasks (P19-2)

Copied to clipboard

Challenge: Existing spelling correction systems are far from perfect for noise-sensitive texts . a new way to handle noise is to make models robust to noise.
Approach: They propose a robust to noise word embeddings model which outperforms existing models in different tasks.
Outcome: The proposed model outperforms existing models in three downstream tasks and shows improvements in noise robustness over existing models.
A Stochastic Decoder for Neural Machine Translation (P18-1)

Copied to clipboard

Challenge: Neural machine translation models do not account for local lexical and syntactic variation in parallel corpora.
Approach: They propose a deep generative model of machine translation which incorporates a chain of latent variables to account for local lexical and syntactic variation in parallel corpora.
Outcome: The proposed model consistently improves over strong baselines on several different language pairs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations