Challenge: Existing multilingual NMT approaches do not utilize the abundance of monolingual data, especially in low-resource languages.
Approach: They propose to combine monolingual data with self-supervision to pre-train translation models and fine-tune on small amounts of supervised data.
Outcome: The proposed approach improves translation quality of low-resource languages and zero-shot translation quality.

Similar Papers

Multilingual Neural Machine Translation (2020.coling-tutorials)

Copied to clipboard

Challenge: In this tutorial, we will cover the latest advances in NMT to enhance low-resource translation.
Approach: They will cover the latest advances in NMT approaches that leverage multilingualism . they will focus on topics such as language divergence, transfer learning and pivoting .
Outcome: This tutorial will cover the latest advances in NMT to enhance low-resource translation models.
Cross-lingual Supervision Improves Unsupervised Neural Machine Translation (2021.naacl-industry)

Copied to clipboard

Challenge: Existing models that use only monolingual data have not been fully duplicated in the vast majority of language pairs, especially for zero-source languages.
Approach: They propose to leverage the corpus from En-Fr and En-De to collectively train the translation from one language into many languages under one model.
Outcome: The proposed model significantly improves translation quality with a big margin in the benchmark unsupervised translation tasks and achieves comparable performance to supervised NMT.
Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to improve multilingual neural machine translation (NMT) are weak, and lack robustness to support language pairs with varying typological characteristics.
Approach: They propose to deepen NMT models to support language pairs with varying typological characteristics by random online backtranslation.
Outcome: The proposed approach narrows the performance gap with bilingual models and improves zero-shot performance by 10 BLEU, approaching conventional pivot-based methods.
Self-Supervised Neural Machine Translation (P19-1)

Copied to clipboard

Challenge: Neural machine translation (NMT) methods relied on the availability of high-quality parallel corpora.
Approach: They propose a method where an emergent NMT system is used for selecting training data and learning internal NMT representations.
Outcome: The proposed method achieves BLEU scores of 29.21 (en2fr) and 27.36 (fr2en) on newstest2014 using English and French Wikipedia data for training.
Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel Data (2021.acl-long)

Copied to clipboard

Challenge: linguistic overlap between low-resource languages and high-resourced languages is a major obstacle for training high-quality machine translation systems.
Approach: They exploit linguistic overlap to facilitate translation to and from low-resource languages . they use monolingual data and parallel data in related high-resourced languages based on their method .
Outcome: The proposed method significantly improves translation into low-resource language compared to baselines on 7 languages from three different language families.
Target Conditioned Sampling: Optimizing Data Selection for Multilingual Neural Machine Translation (P19-1)

Copied to clipboard

Challenge: Existing studies show that training on a single related language is more effective than using all data.
Approach: They propose an efficient algorithm that first samples a target sentence, and then conditionally samples its source sentence.
Outcome: The proposed algorithm brings significant gains on three of four languages with minimal training overhead.
Self-Training Sampling with Monolingual Data Uncertainty for Neural Machine Translation (2021.acl-long)

Copied to clipboard

Challenge: Experimental results show that enhancing the learning on uncertain monolingual sentences improves the translation quality of high-uncertainty sentences and also benefits the prediction of low-frequency words at the target side.
Approach: They propose to use monolingual data to augment model training with synthetic parallel data by selecting the most informative monolingual sentences to complement the parallel data.
Outcome: The proposed approach improves the performance of natural language models by selecting the most informative monolingual sentences.
Handling Syntactic Divergence in Low-resource Machine Translation (D19-1)

Copied to clipboard

Challenge: Existing approaches to neural machine translation (NMT) are dependent on limited parallel data, and can be difficult to use for many language pairs.
Approach: They propose a method where target-language sentences are re-ordered to match the order of the source and used as an additional source of training-time supervision.
Outcome: The proposed method improves on simulated low-resource Japanese-to-English and real low-demand Uyghur-to English scenarios.
Revisiting Low-Resource Neural Machine Translation: A Case Study (P19-1)

Copied to clipboard

Challenge: Recent research has shown that neural machine translation models are highly data-inefficient and underperform phrase-based statistical machine translation (PBSMT) in low-resource settings.
Approach: They propose to use auxiliary data to train low-resource neural machine translation systems without auxiliary monolingual or multilingual data.
Outcome: The proposed methods outperform PBSMT and other statistical machine translation models in Korean–English with minimal data.
Meta-Learning for Low-Resource Neural Machine Translation (D18-1)

Copied to clipboard

Challenge: In this paper, we propose to extend the recently introduced model-agnostic meta-learning algorithm for low-resource neural machine translation (NMT).
Approach: They propose to extend the recently introduced meta-learning algorithm for low-resource neural machine translation (NMT) they frame low-Resource translation as a meta- learning problem where we learn to adapt to low-REsource languages based on multilingual high-resourced language tasks.
Outcome: The proposed meta-learning algorithm outperforms the multilingual, transfer learning based approach and can train a competitive NMT system with only a fraction of training examples.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations