Generalized Data Augmentation for Low-Resource Translation (P19-1)

Copied to clipboard

Challenge: Low-resource language pairs with a lack of parallel data pose challenges for machine translation . data augmentation using monolingual data is an effective way to alleviate the problem .
Approach: They propose a general framework for data augmentation for low-resource machine translation using monolingual data and a related high-resourced language.
Outcome: The proposed method improves translation quality by 1.5 to 8 BLEU points under extreme low-resource settings compared to baselines.

Similar Papers

Rethinking Data Augmentation for Low-Resource Neural Machine Translation: A Multi-Task Learning Approach (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to generating additional parallel sentences are aimed at expanding the support of the empirical data distribution by generating new sentence pairs that contain infrequent words.
Approach: They propose to use data augmentation techniques to generate additional parallel sentences by reversing the order of the target sentence to produce unfluent target sentences.
Outcome: The proposed approach improves on six low-resource translation tasks and the baseline and over DA methods.
A systematic comparison of methods for low-resource dependency parsing on genuinely low-resource languages (D19-1)

Copied to clipboard

Challenge: Large annotated treebanks are available for only a tiny fraction of the world's languages, and there is a wealth of literature on strategies for parsing with few resources.
Approach: They propose three strategies for improving low-resource parsers: data augmentation, cross-lingual training, and transliteration.
Outcome: The proposed methods improve low-resource parsers by using data augmentation, cross-lingual training, and transliteration.
Language Model Priors and Data Augmentation Strategies for Low-resource Machine Translation: A Case Study Using Finnish to Northern Sámi (2024.findings-acl)

Copied to clipboard

Challenge: a new study examines the use of monolingual data for improving low-resource machine translation.
Approach: They investigate ways of using monolingual data for improving low-resource machine translation.
Outcome: The proposed model can perform better on the target-side data without augmentation of parallel data.
Small Data, Big Impact: Leveraging Minimal Data for Effective Machine Translation (2023.acl-long)

Copied to clipboard

Challenge: Existing datasets are not economical to create large-scale datasets, but for low-resource languages, a few thousand professionally translated sentence pairs can be useful.
Approach: They propose to use a dataset to train machine translation models on pre-existing and synthetic data to augment them with millions of sentences through backtranslation.
Outcome: The proposed model can cover hundreds of languages with high quality training data even when smaller but lower quality datasets are used.
Data Augmentation via Subtree Swapping for Dependency Parsing of Low-Resource Languages (2020.coling-main)

Copied to clipboard

Challenge: Lack of annotated training data is a big issue for building reliable NLP systems for most of the world’s languages.
Approach: They propose a method to swap subtrees between annotated sentences while enforcing strong constraints on those trees to ensure maximum grammaticality of the new sentences.
Outcome: The proposed method outperforms previous methods using the same inputs and using low-resource languages.
Grammar-based Data Augmentation for Low-Resource Languages: The Case of Guarani-Spanish Neural Machine Translation (2024.naacl-long)

Copied to clipboard

Challenge: Low-resource languages suffer from a vicious circle: data is needed to build tools, but available text is scarce.
Approach: They propose to use a grammar-based system to generate Spanish text and syntactically transfer it to Guarani to boost its performance.
Outcome: The proposed system outperforms existing models by pretraining models with synthetic text.
Few-Shot Learning Translation from New Languages (2025.emnlp-main)

Copied to clipboard

Challenge: Recent work shows strong transfer learning capability to unseen languages in sequence-to-sequence neural networks . current transfer learning methods require much less downstream task data than would otherwise be required.
Approach: They first train word embeddings models on varying amounts of data and plug them into a machine translation model.
Outcome: The proposed model can learn Flores with only 500 parallel sentences and 31,250 sentences of monolingual data, and it can exceed 15 BLEU on unseen languages.
Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel Data (2021.acl-long)

Copied to clipboard

Challenge: linguistic overlap between low-resource languages and high-resourced languages is a major obstacle for training high-quality machine translation systems.
Approach: They exploit linguistic overlap to facilitate translation to and from low-resource languages . they use monolingual data and parallel data in related high-resourced languages based on their method .
Outcome: The proposed method significantly improves translation into low-resource language compared to baselines on 7 languages from three different language families.
Handling Syntactic Divergence in Low-resource Machine Translation (D19-1)

Copied to clipboard

Challenge: Existing approaches to neural machine translation (NMT) are dependent on limited parallel data, and can be difficult to use for many language pairs.
Approach: They propose a method where target-language sentences are re-ordered to match the order of the source and used as an additional source of training-time supervision.
Outcome: The proposed method improves on simulated low-resource Japanese-to-English and real low-demand Uyghur-to English scenarios.
Getting More Data for Low-resource Morphological Inflection: Language Models and Data Augmentation (2020.lrec-1)

Copied to clipboard

Challenge: Morphological inflection is the process that generates the word form given its lexeme and morphological properties.
Approach: They propose to use language models and data augmentation to improve morphological inflection without annotating more data.
Outcome: The proposed model improves by 1.5% with the langauge model and by 9% with the data augmentation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations