Challenge: Prior work has shown that translating from multiple source languages improves translation quality.
Approach: They propose to exploit multiple source sentences from auxiliary languages to improve multilingual translation in a more common scenario by using synthetic multi-source corpora.
Outcome: Extensive experiments on Chinese/English-Japanese and a large-scale multilingual translation benchmark show that the proposed model outperforms the baseline model significantly by +4.0 BLEU.

Similar Papers

Improving a Multi-Source Neural Machine Translation Model with Corpus Extension for Low-Resource Languages (L18-1)

Copied to clipboard

Challenge: In machine translation, we often try to collect resources to improve performance.
Approach: They propose to use synthetic methods to extend low-resource corpus to create target sentences using synthetic methods.
Outcome: The proposed method improves translation performance for low-resource language pairs.
Improving Multilingual Neural Machine Translation by Utilizing Semantic and Linguistic Features (2024.findings-acl)

Copied to clipboard

Challenge: Existing models do not differentiate between semantic and linguistic features, resulting in the entanglement of knowledge and linguistics within the model.
Approach: They propose to exploit both semantic and linguistic features to enhance multilingual translation by disentangling encoder representations and integrating low-level linguistic encoders.
Outcome: The proposed model improves zero-shot translation while maintaining performance in supervised translation on multilingual datasets.
An Analysis of Massively Multilingual Neural Machine Translation for Low-Resource Languages (2020.lrec-1)

Copied to clipboard

Challenge: In this study, we explore massively multilingual low-resource neural machine translation.
Approach: They propose to use Bible translations to train models with up to 1,107 source languages and create multilingual corpora varying the number and relatedness of source languages.
Outcome: The proposed approach is highly language-specific and can be tailored to the source language and its typology.
Improving Language Model Integration for Neural Machine Translation (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to integrate external language models into machine translation systems have been based on the assumption that the external model learns an implicit target-side language model at decoding time.
Approach: They transfer this concept to the task of machine translation and compare it with the most prominent way of including additional monolingual data - namely back-translation.
Outcome: The proposed approach outperforms the most prominent way of including additional monolingual data, namely back-translation.
Multi-Source Syntactic Neural Machine Translation (D18-1)

Copied to clipboard

Challenge: Existing approaches to integrate source syntax into neural machine translations use linearized parses.
Approach: They propose a linearized parsed neural machine translation technique that integrates source syntax into neural machine learning.
Outcome: The proposed model improves over seq2seq and parsed baselines by over 1 BLEU on the WMT17 English-German task.
Multilingual Neural Machine Translation (2020.coling-tutorials)

Copied to clipboard

Challenge: In this tutorial, we will cover the latest advances in NMT to enhance low-resource translation.
Approach: They will cover the latest advances in NMT approaches that leverage multilingualism . they will focus on topics such as language divergence, transfer learning and pivoting .
Outcome: This tutorial will cover the latest advances in NMT to enhance low-resource translation models.
Neural Machine Translation for Bilingually Scarce Scenarios: a Deep Multi-Task Learning Approach (N18-1)

Copied to clipboard

Challenge: Neural machine translation requires large amount of parallel training text to learn a reasonable quality translation model.
Approach: They propose a multi-task learning approach that leverages monolingual linguistic resources in the source side of a machine translation task.
Outcome: The proposed approach is effective on three translation tasks: English-to-French, English- to-Farsi, and English-à-Vietnamese.
Multilingual Unsupervised Neural Machine Translation with Denoising Adapters (2021.emnlp-main)

Copied to clipboard

Challenge: Multilingual unsupervised machine translation is a computationally expensive and hard to tune approach . auxiliary parallel data is used to train translation systems from monolingual data .
Approach: They propose to use auxiliary parallel language pairs to train unsupervised machine translations . they propose to add auxiliary languages to pre-trained mBART-50 models with denoising adapters .
Outcome: The proposed approach is on-par with back-translation and allows adding unseen languages incrementally.
Efficient Inference for Multilingual Neural Machine Translation (2021.emnlp-main)

Copied to clipboard

Challenge: Multilingual NMT is an attractive solution for production, but to match bilingual quality, it comes at the cost of larger and slower models.
Approach: They propose to use a shallow decoder with vocabulary filtering to speed up inference . they validate their findings with BLEU and chrF on 380 language pairs .
Outcome: The proposed approach can be used in two 20-language multi-parallel settings.
Massively Multilingual Neural Machine Translation (N19-1)

Copied to clipboard

Challenge: Multilingual Neural Machine Translation models support translation from multiple source languages into multiple target languages.
Approach: They perform extensive experiments in training massively multilingual NMT models involving up to 103 distinct languages and 204 translation directions simultaneously.
Outcome: The proposed model outperforms the state-of-the-art in low resource settings while supporting up to 59 languages in 116 translation directions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations