Prediction Difference Regularization against Perturbation for Neural Machine Translation (2022.acl-long)
Copied to clipboard
| Challenge: | Existing methods for regularizing input perturbation are limited by under-fitting of training data. |
| Approach: | They propose a method that can reduce over-fitting and under-fitting at the same time. |
| Outcome: | The proposed method can reduce over-fitting and under-fitturing while making the model less sensitive to small input changes and more robust to under-perturbed training data. |
Similar Papers
Effective Adversarial Regularization for Neural Machine Translation (P19-1)
Copied to clipboard
| Challenge: | Existing (small) perturbations that induce a critical prediction error in machine learning models are often referred to as adversarial examples. |
| Approach: | They propose to use adversarial perturbations to regularize text classification tasks by adding adversarials to a typical NMT model structure. |
| Outcome: | The proposed method significantly improves performance of NMT models, such as LSTM-based and Transformer-based models. |
Bridging the Gap between Training and Inference for Neural Machine Translation (P19-1)
Copied to clipboard
| Challenge: | Neural Machine Translation generates target words sequentially while at inference it has to generate the entire sequence from scratch. |
| Approach: | They propose to use ground truth and inference to generate target words sequentially while at inference it has to generate the entire sequence from scratch. |
| Outcome: | Experiments on Chinese->English and WMT’14 English->German translation tasks show that the proposed model can achieve significant improvements on multiple datasets. |
Evaluating Robustness to Input Perturbations for Neural Machine Translation (2020.acl-main)
Copied to clipboard
| Challenge: | Recent work has shown that Neural Machine Translation models are brittle to small perturbations in the input. |
| Approach: | They propose to use subword regularization to measure the relative degradation and changes in translation when perturbations are added to the input. |
| Outcome: | The proposed measures show that the models are more robust to perturbations when subword regularization methods are used. |
Revisiting Negation in Neural Machine Translation (2021.tacl-1)
Copied to clipboard
| Challenge: | Negation is an important linguistic phenomenon in machine translation, as errors in translating negation may change the meaning of source sentences completely. |
| Approach: | They evaluate the translation of negation in English–German (EN–DE) and English– Chinese (EN-ZH) . they find that NMT models can distinguish negation and non-negation tokens very well and encode a lot of information about negation . |
| Outcome: | The accuracy of manual evaluation in ENDE, DEEN, ENZH, and ZHEN is 95.7%, 94.8%, 93.4%, and 91.7% respectively. |
Beyond Noise: Mitigating the Impact of Fine-grained Semantic Divergences on Neural Machine Translation (2021.acl-long)
Copied to clipboard
| Challenge: | Prior work treats all types of mismatches between source and target as noise . Consequently, it remains unclear how noisy parallel training samples impact NMT training. |
| Approach: | They propose a divergent-aware NMT framework that uses factors to help NMT recover from the degradation caused by naturally occurring divergences. |
| Outcome: | The proposed framework improves translation quality and model calibration on EN-FR tasks. |
Revisiting Low-Resource Neural Machine Translation: A Case Study (P19-1)
Copied to clipboard
| Challenge: | Recent research has shown that neural machine translation models are highly data-inefficient and underperform phrase-based statistical machine translation (PBSMT) in low-resource settings. |
| Approach: | They propose to use auxiliary data to train low-resource neural machine translation systems without auxiliary monolingual or multilingual data. |
| Outcome: | The proposed methods outperform PBSMT and other statistical machine translation models in Korean–English with minimal data. |
On the Inference Calibration of Neural Machine Translation (2020.acl-main)
Copied to clipboard
| Challenge: | Existing studies show that NMT models trained with label smoothing are well-calibrated on ground-truth training data, but miscalibration remains a challenge during inference due to the discrepancy between training and inference. |
| Approach: | They propose a graduated label smoothing method that can improve inference calibration and translation performance. |
| Outcome: | The proposed method improves both inference calibration and translation performance. |
Token Drop mechanism for Neural Machine Translation (2020.coling-main)
Copied to clipboard
| Challenge: | Neural machine translation models are vulnerable to unfamiliar inputs. |
| Approach: | They propose to drop tokens of the input sentences to improve generalization and avoid overfitting for the NMT model. |
| Outcome: | The proposed approach improves on Chinese-English and English-Romanian benchmarks and achieves significant performance improvements over baselines. |
On the Sparsity of Neural Machine Translation Models (2020.emnlp-main)
Copied to clipboard
| Challenge: | Modern neural machine translation models employ a large number of parameters, which leads to serious over-parameterization. |
| Approach: | They propose to prune parameters to improve the model by +0.8 BLEU points and to reallocate them to enhance the ability of modeling low-level lexical information. |
| Outcome: | The pruned parameters improve the model by +0.8 BLEU points and the rejuvenated parameters enhance the ability to model low-level lexical information. |
Bi-Directional Differentiable Input Reconstruction for Low-Resource Neural Machine Translation (N19-1)
Copied to clipboard
| Challenge: | Existing work has addressed this problem by leveraging monolingual or multilingual data. |
| Approach: | They propose to introduce a differentiable reconstruction loss for neural machine translation to exploit the limited amounts of parallel text available in low-resource settings. |
| Outcome: | The proposed approach achieves small but consistent BLEU improvements on four language pairs in both translation directions and outperforms an alternative differentiable reconstruction strategy based on hidden states. |