Challenge: Low-resource African languages pose unique challenges for natural language processing (NLG) We demonstrate the effectiveness of Cheetah through comprehensive evaluations across six generation downstream tasks.
Approach: They develop a multilingual NLG language model for African languages called Cheetah . they demonstrate that Cheethah outperforms other models in six tasks .
Outcome: The proposed model outperforms other models in five of six generation tasks.

Similar Papers

Toucan: Many-to-Many Translation for 150 African Language Pairs (2024.findings-acl)

Copied to clipboard

Challenge: We introduce two language models with 1.2 billion and 3.7 billion parameters to improve Machine Translation (MT) for low-resource languages.
Approach: They propose a set of tools to improve Machine Translation (MT) for low-resource languages with a focus on African languages.
Outcome: The proposed model outperforms existing models on MT for African languages and improves translation evaluation metrics for 1K languages including African languages.
Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead (2025.emnlp-main)

Copied to clipboard

Challenge: African languages are often left behind in state-of-the-art natural language processing systems and large language models.
Approach: They analyze 884 research papers on NLP for African languages published over past five years . they identify key trends shaping the field and outline promising directions .
Outcome: The findings identify key trends shaping the field and outline promising directions . the authors analyze 884 research papers on NLP for African languages published over the past five years .
AfroBench: How Good are Large Language Models on African Languages? (2025.findings-acl)

Copied to clipboard

Challenge: Large-scale multilingual evaluations often include only a handful of African languages due to the scarcity of high-quality data and the limited discoverability of existing datasets.
Approach: They propose a multi-task benchmark to evaluate the performance of LLMs across 64 African languages, 15 tasks and 22 datasets.
Outcome: The proposed benchmark compares LLMs across 64 African languages, 15 tasks and 22 datasets.
The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLP (2026.acl-long)

Copied to clipboard

Challenge: Among the approximately 7,000 languages spoken globally, fewer than 20 receive substantial attention in NLP research.
Approach: They propose to use African multi-modal speech and text data to validate African multimodal models and validate them on targeted language data.
Outcome: The African Languages Lab's results show that the proposed model outperforms untrained models in 31 languages and a 1B-parameter model beats the commercial system in Yoruba and Twi.
SERENGETI: Massively Multilingual Language Models for Africa (2023.findings-acl)

Copied to clipboard

Challenge: Pretrained models acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning.
Approach: They develop a set of massively multilingual language models that covers 517 African languages and language varieties.
Outcome: The proposed models outperform 4 models that cover 4-23 African languages on eight natural language understanding tasks, achieving 82.27 average F_1.
IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are limited to a few high-resource languages . many low-resourced languages are evaluated only on basic text classification tasks .
Approach: They propose to use IrokoBench to evaluate 17 low-resource African languages . they use human-translated benchmark datasets to evaluate zero-shot, few-shot and translate-test settings .
Outcome: The proposed model performs well in English and French, but the highest performing model perform poorly in proprietary models.
AfroMT: Pretraining Strategies and Reproducible Benchmarks for Translation of 8 African Languages (2021.emnlp-main)

Copied to clipboard

Challenge: Existing reproducible benchmarks for machine translation are limited to high-resource or well-represented languages.
Approach: They propose to use AfroMT to develop a reproducible machine translation benchmark for eight widely spoken African languages and a suite of analysis tools to take into account their unique properties.
Outcome: The proposed benchmarks show significant improvements when pretraining on 11 languages, with gains of up to 2 BLEU points over strong baselines.
Towards Afrocentric NLP for African Languages: Where We Are and Where We Can Go (2022.acl-long)

Copied to clipboard

Challenge: ACL 2022 special Theme on "Language Diversity: from Low Resource to Endangered Languages" focuses on linguistic and sociopolitical challenges facing development of NLP technologies for African languages .
Approach: They propose a typological framework for linguistic and sociopolitical challenges for NLP in African languages.
Outcome: The main objective of this study is to motivate and advocate for an Afrocentric approach to technology development.
Multilingual Generation in Abstractive Summarization: A Comparative Study (2024.lrec-main)

Copied to clipboard

Challenge: Existing models for multilingual generation lack thorough analysis due to extensive linguistic diversity.
Approach: They propose to classify multilingual generation methodologies into three categories based on their underlying modeling principles . they introduce an automatic metric to mitigate spurious correlations associated with language mixing .
Outcome: The proposed model improves in high-resource, low-resourced, and zero-shot scenarios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations