Few-Shot Learning Translation from New Languages (2025.emnlp-main)

Copied to clipboard

Challenge: Recent work shows strong transfer learning capability to unseen languages in sequence-to-sequence neural networks . current transfer learning methods require much less downstream task data than would otherwise be required.
Approach: They first train word embeddings models on varying amounts of data and plug them into a machine translation model.
Outcome: The proposed model can learn Flores with only 500 parallel sentences and 31,250 sentences of monolingual data, and it can exceed 15 BLEU on unseen languages.

Similar Papers

An Analysis of Massively Multilingual Neural Machine Translation for Low-Resource Languages (2020.lrec-1)

Copied to clipboard

Challenge: In this study, we explore massively multilingual low-resource neural machine translation.
Approach: They propose to use Bible translations to train models with up to 1,107 source languages and create multilingual corpora varying the number and relatedness of source languages.
Outcome: The proposed approach is highly language-specific and can be tailored to the source language and its typology.
Small Data, Big Impact: Leveraging Minimal Data for Effective Machine Translation (2023.acl-long)

Copied to clipboard

Challenge: Existing datasets are not economical to create large-scale datasets, but for low-resource languages, a few thousand professionally translated sentence pairs can be useful.
Approach: They propose to use a dataset to train machine translation models on pre-existing and synthetic data to augment them with millions of sentences through backtranslation.
Outcome: The proposed model can cover hundreds of languages with high quality training data even when smaller but lower quality datasets are used.
Universal Neural Machine Translation for Extremely Low Resource Languages (N18-1)

Copied to clipboard

Challenge: a novel multilingual approach to machine translation is proposed for low resource languages . the proposed approach can achieve 23 BLEU on Romanian-English WMT2016 using a tiny parallel corpus of 6k sentences compared to the 18 BLUE of strong baseline system .
Approach: They propose a transfer-learning approach to share lexical and sentence representations across multiple source languages into one target language.
Outcome: The proposed approach achieves 23 BLEU on Romanian-English WMT2016 using a tiny parallel corpus of 6k sentences compared to the 18 BLUE of strong baseline system which uses multi-lingual training and back-translation.
Multilingual unsupervised sequence segmentation transfers to extremely low-resource languages (2022.acl-long)

Copied to clipboard

Challenge: Unsupervised sequence segmentation is a key component of low-resource languages where there is little or no gold-standard data on which to train supervised models.
Approach: They propose to pre-train a Masked Segmental Language Model multilingually to achieve unsupervised segmentation performance in extremely low-resource languages.
Outcome: The proposed model outperforms a monolingual model and a pre-trained model on Quechua in 6/10 settings.
Disentangling Pretrained Representation to Leverage Low-Resource Languages in Multilingual Machine Translation (2024.lrec-main)

Copied to clipboard

Challenge: Multilingual neural machine translation requires an enormous dataset, leaving the low-resource language (LRL) underdeveloped.
Approach: They evaluated five languages using a parallel corpus of 1,000 instances each and found a zero-shot improvement of 7.4 from the baseline score of 7.1 to a score of 15.5 at best.
Outcome: The proposed model improves performance in the linguistically diverse country of Indonesia by 7.4 from baseline score of 7.1 to 15.5 at best.
In Neural Machine Translation, What Does Transfer Learning Transfer? (2020.acl-main)

Copied to clipboard

Challenge: a recent study found that word embeddings are not necessary for transfer learning.
Approach: They perform several ablation studies that limit information transfer and measure the quality impact across three language pairs to gain a black-box understanding of transfer learning.
Outcome: The proposed method can eliminate the need for a warm-up phase when training transformer models in high resource language pairs.
A Simple and Effective Method to Improve Zero-Shot Cross-Lingual Transfer Learning (2022.coling-1)

Copied to clipboard

Challenge: Existing zero-shot cross-lingual transfer methods rely on parallel corpora or bilingual dictionaries . however, its effect is limited by the gap between embedding clusters of different languages .
Approach: They propose Embedding-Push, Attention-Pull, and Robust targets to transfer English embeddings to virtual multilingual embedders without semantic loss.
Outcome: Experimental results show that the proposed method outperforms existing methods on cross-lingual tasks and can achieve a better multilingual alignment.
When Being Unseen from mBERT is just the Beginning: Handling New Languages With Multilingual Language Models (2021.naacl-main)

Copied to clipboard

Challenge: Language models are a new standard to build state-of-the-art NLP systems.
Approach: They compare multilingual and monolingual models on unseen languages . they show that some languages benefit from transfer learning whereas others don't .
Outcome: The proposed model behaves in multiple ways on unseen languages, while others fail to transfer . the results provide a promising direction towards making multilingual models useful for a new set of unseense languages.
Low-resource Neural Machine Translation with Cross-modal Alignment (2022.emnlp-main)

Copied to clipboard

Challenge: Existing neural machine translation techniques rely on large monolingual corpus, which is costly for some low-resource languages.
Approach: They propose a cross-modal contrastive learning method to learn a shared space for all languages by additional visual modality.
Outcome: The proposed method can learn cross-modal and cross-lingual alignment with small amount of image-text pairs and achieves significant improvements over the text-only baseline.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations