Bayelemabaga: Creating Resources for Bambara NLP (2025.naacl-long)

Copied to clipboard

Challenge: a lack of well-structured multilingual datasets remains a challenge for machine translation in under-resource languages.
Approach: They propose to create a multilingual dataset for machine translation in the Bambara language, the vehicular language of Mali.
Outcome: The proposed dataset is the most extensive curated multilingual dataset for machine translation in the Bambara language, the vehicular language of Mali.

Similar Papers

Kumatigi: Quality-Driven Data Augmentation for Low-Resource Machine Translation (2026.findings-acl)

Copied to clipboard

Challenge: Neural machine translation for extremely low-resource languages faces compounding challenges: limited parallel data, orthographic inconsistency, and inconsistent metadata for principled training.
Approach: They propose a quality-annotated French-Bambara corpus combining systematic curation with data augmentation strategies tailored to Bambaran.
Outcome: The proposed framework achieves up to +3–4 BLEU over strong baselines.
Toucan: Many-to-Many Translation for 150 African Language Pairs (2024.findings-acl)

Copied to clipboard

Challenge: We introduce two language models with 1.2 billion and 3.7 billion parameters to improve Machine Translation (MT) for low-resource languages.
Approach: They propose a set of tools to improve Machine Translation (MT) for low-resource languages with a focus on African languages.
Outcome: The proposed model outperforms existing models on MT for African languages and improves translation evaluation metrics for 1K languages including African languages.
TaTA: A Multilingual Table-to-Text Dataset for African Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing data-to-text generation datasets are limited to English and a small number of other languages.
Approach: They create the first large multilingual table-to-text dataset with a focus on African languages.
Outcome: The proposed dataset includes 8,700 examples in nine languages including four African languages and a zero-shot test language.
AfroMT: Pretraining Strategies and Reproducible Benchmarks for Translation of 8 African Languages (2021.emnlp-main)

Copied to clipboard

Challenge: Existing reproducible benchmarks for machine translation are limited to high-resource or well-represented languages.
Approach: They propose to use AfroMT to develop a reproducible machine translation benchmark for eight widely spoken African languages and a suite of analysis tools to take into account their unique properties.
Outcome: The proposed benchmarks show significant improvements when pretraining on 11 languages, with gains of up to 2 BLEU points over strong baselines.
AfriMMT-EA: Multi-domain Machine Translation for Low-Resource East African Languages (2026.findings-eacl)

Copied to clipboard

Challenge: Recent advances in open-source large language models have demonstrated strong multilingual capabilities through data-efficient adaptation strategies.
Approach: They propose to use AfriMMT-EA to refine two multilingual versions of Gemma-3 to better understand the region's linguistic and cultural diversity.
Outcome: The proposed datasets comprise 54 local languages across five East African countries.
Massive vs. Curated Embeddings for Low-Resourced Languages: the Case of Yorùbá and Twi (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that word embeddings can be useful for training downstream natural language processing tasks.
Approach: They compare word embeddings obtained by word embeds from curated corpora with a language-dependent processing.
Outcome: The proposed model compares word embeddings with word embeds from curated corpora and a language-dependent processing on two African languages.
AFRIDOC-MT: Document-level MT Corpus for African Languages (2025.emnlp-main)

Copied to clipboard

Challenge: AFRIDOC-MT is a document-level multi-parallel translation dataset covering five languages . AFRITIC-MT models perform better on sentences than general-purpose LLMs .
Approach: They propose a document-level multi-parallel translation dataset covering English and five African languages.
Outcome: The proposed dataset covers 334 health and 271 information technology news documents . it shows that NLLB-200 achieves the best average performance among standard models .
Better Quality Pre-training Data and T5 Models for African Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Existing web crawls have demonstrated quality issues for low-resource languages . Existing pretraining corpora have numerous quality issues .
Approach: They propose to audit existing pretraining corpora to understand and rectify quality issues . they pretrain a new T5-based model and evaluate its performance on multiple tasks .
Outcome: The proposed model outperforms existing pretrained models on four NLP tasks.
Cheetah: Natural Language Generation for 517 African Languages (2024.acl-long)

Copied to clipboard

Challenge: Low-resource African languages pose unique challenges for natural language processing (NLG) We demonstrate the effectiveness of Cheetah through comprehensive evaluations across six generation downstream tasks.
Approach: They develop a multilingual NLG language model for African languages called Cheetah . they demonstrate that Cheethah outperforms other models in six tasks .
Outcome: The proposed model outperforms other models in five of six generation tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations