Challenge: Creole languages are used in much of Latin America, Africa and the Caribbean . a large multilingual bitext like ours has potential to build the best yet or first ever MT models for many languages .
Approach: They present the largest cumulative dataset to date for Creole language MT . they provide MT models supporting all 41 Creoles in 172 translation directions .
Outcome: The proposed model outperforms a genre-specific Creole MT model on its own benchmark for 23 of 34 translation directions.

Similar Papers

AfriMMT-EA: Multi-domain Machine Translation for Low-Resource East African Languages (2026.findings-eacl)

Copied to clipboard

Challenge: Recent advances in open-source large language models have demonstrated strong multilingual capabilities through data-efficient adaptation strategies.
Approach: They propose to use AfriMMT-EA to refine two multilingual versions of Gemma-3 to better understand the region's linguistic and cultural diversity.
Outcome: The proposed datasets comprise 54 local languages across five East African countries.
AFRIDOC-MT: Document-level MT Corpus for African Languages (2025.emnlp-main)

Copied to clipboard

Challenge: AFRIDOC-MT is a document-level multi-parallel translation dataset covering five languages . AFRITIC-MT models perform better on sentences than general-purpose LLMs .
Approach: They propose a document-level multi-parallel translation dataset covering English and five African languages.
Outcome: The proposed dataset covers 334 health and 271 information technology news documents . it shows that NLLB-200 achieves the best average performance among standard models .
KreolMorisienMT: A Dataset for Mauritian Creole Machine Translation (2022.findings-aacl)

Copied to clipboard

Challenge: Mauritian Creole is a French-based creole and a lingua franca of the Republic of Mauritius.
Approach: They describe a dataset for benchmarking machine translation quality of Mauritian Creole.
Outcome: The proposed dataset compares KreolMorisienMT with existing models and human evaluation reveals the systems’ high translation quality.
SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages? (2025.emnlp-main)

Copied to clipboard

Challenge: Existing metrics for machine translation quality for under-resourced African languages suffer from limited language coverage and poor performance in low-resource settings.
Approach: They propose a large-scale human-annotated machine translation evaluation dataset . they use a reference-based and reference-free evaluation model to compare MT quality .
Outcome: The proposed models outperform AfriCOMET and the strongest LLM on low-resource languages.
Machine Translation into Low-resource Language Varieties (2021.acl-short)

Copied to clipboard

Challenge: Current machine translation systems generate a "standard" target language, but many languages have multiple varieties that are different from the standard language.
Approach: They propose a framework to rapidly adapt machine translation systems to generate different target varieties . they propose to use no parallel data to generate languages close to, but different from, the standard target language .
Outcome: The proposed model improves on a system that generates Ukrainian and Belarusian in two languages with no parallel data.
Benchmarking Neural and Statistical Machine Translation on Low-Resource African Languages (2020.lrec-1)

Copied to clipboard

Challenge: a recent study has focused on languages where large amounts of resources are available.
Approach: They benchmark state of the art statistical and neural machine translation systems on Somali and Swahili languages . they find that statistical machine translation and neural translation can perform similarly in low-resource scenarios .
Outcome: The results show that statistical machine translation and neural machine translation perform similarly in low-resource scenarios.
Toucan: Many-to-Many Translation for 150 African Language Pairs (2024.findings-acl)

Copied to clipboard

Challenge: We introduce two language models with 1.2 billion and 3.7 billion parameters to improve Machine Translation (MT) for low-resource languages.
Approach: They propose a set of tools to improve Machine Translation (MT) for low-resource languages with a focus on African languages.
Outcome: The proposed model outperforms existing models on MT for African languages and improves translation evaluation metrics for 1K languages including African languages.
Many-to-English Machine Translation Tools, Data, and Pretrained Models (2021.acl-demo)

Copied to clipboard

Challenge: Commercial translation systems support only one hundred languages or fewer . commercial translation systems do not make these models available for transfer to low resource languages .
Approach: They propose a multilingual neural machine translation model that can translate from 500 source languages to English.
Outcome: The proposed model can translate from 500 source languages to English, or be used as a parent model for low-resource languages.
Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown continuously improving multilingual capabilities.
Approach: They evaluate the ability of open LLMs to handle multilingual machine translation tasks using a parallel-first monolingual-second data mixing strategy.
Outcome: The proposed model outperforms state-of-the-art models and achieves competitive performance with Google Translate and GPT-4-turbo.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations