Challenge: AfroLID is a neural LID toolkit for 517 African languages and varieties.
Approach: They propose to exploit a multi-domain web dataset manually curated from across 14 language families utilizing five orthographic systems to exploit AfroLID.
Outcome: The proposed tool outperforms existing tools on the acutely under-served Twitter domain.

Similar Papers

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data (2026.acl-long)

Copied to clipboard

Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera, Sara Hincapié Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, Shamsuddeen Hassan Muhammad, Atnafu Lambebo Tonja, Hend Al-Khalifa, Nadia Ghezaiel Hammouda, Verrah Akinyi Otiende, Tack Hwa Wong, Jakhongir Saydaliev, Melika Nobakhtian, Muhammad Ravi Shulthan Habibi, Chalamalasetti Kranti, Carol Muchemi, Khang Nguyen, Faisal Muhammad Adam, Luis Frentzen Salim, Reem Alqifari, Cynthia Jayne Amol, Joseph Marvin Imperial, Ilker Kesen, Ahmad Mustafid, Pavel Stepachev, Leshem Choshen, David Anugraha, Hamada Nayel, Seid Muhie Yimam, Vallerie Alexandra Putra, My Chiffon Nguyen, Azmine Toushik Wasi, Gouthami Vadithya, Rob Van Der Goot, Lanwenn ar C’horr, Karan Dua, Andrew Yates, Mithil Bangera, Yeshil Bangera, Hitesh Laxmichand Patel, Shu Okabe, Fenal Ashokbhai Ilasariya, Dmitry Gaynullin, Genta Indra Winata, Yiyuan Li, Juan Pablo Martínez, Amit Agarwal, Ikhlasul Akmal Hanif, Raia Abu Ahmad, Esther Adenuga, Filbert Aurelian Tjiaranata, Weerayut Buaphet, Michael Anugraha, Sowmya Vajjala, Benjamin L Rice, Azril Hafizi Amirudin, Jesujoba Oluwadara Alabi, Srikant Panda, Yassine Toughrai, Bruhan Kyomuhendo, Daniel Ruffinelli, null Akshata, Manuel Goulão, Ej Zhou, Ingrid Gabriela Franco Ramirez, Cristina Aggazzotti, Konstantin Dobler, Jun Kevin, Quentin Pagès, Nicholas Andrews, Nuhu Ibrahim, Mattes Ruckdeschel, Amr Keleg, Mike Zhang, Casper Rufaro Muziri, Saron Samuel, Sotaro Takeshita, Kun Kerdthaisong, Luca Foppiano, Rasul Dent, Tommaso Green, Ahmad Mustapha Wali, Kamohelo Makaaka, Vicky Feliren, Inshirah Idris, Hande Celikkanat, Abdulhamid Abubakar, Jean Maillard, Benoît Sagot, Thibault Clérice, Kenton Murray, Sarah K. K. Luger
Challenge: Language identification (LID) is a fundamental step in curating multilingual corpora.
Approach: They introduce CommonLID, a community-driven, human-annotated LID benchmark for the web domain, covering 109 languages.
Outcome: The proposed benchmark covers 109 languages and shows that existing evaluations overestimate accuracy for many languages in the web domain.
GlotLID: Language Identification for Low-Resource Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing web-mined datasets for low-resource languages have been useful for low resource NLP.
Approach: They propose a model that identifies 1665 low-resource languages and a new model that is rigorously evaluated and reliable.
Outcome: The proposed model outperforms baselines when balancing F1 and false positive rate (FPR).
AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Africa has the highest linguistic diversity among all continents.
Approach: They introduce a sentiment analysis benchmark that contains >110,000 tweets in 14 African languages . they describe the data collection methodology, annotation process, and challenges .
Outcome: The proposed dataset contains >110,000 tweets in 14 African languages . the tweets were annotated by native speakers and used in the shared task .
The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLP (2026.acl-long)

Copied to clipboard

Challenge: Among the approximately 7,000 languages spoken globally, fewer than 20 receive substantial attention in NLP research.
Approach: They propose to use African multi-modal speech and text data to validate African multimodal models and validate them on targeted language data.
Outcome: The African Languages Lab's results show that the proposed model outperforms untrained models in 31 languages and a 1B-parameter model beats the commercial system in Yoruba and Twi.
MasakhaNER: Named Entity Recognition for African Languages (2021.tacl-1)

Copied to clipboard

Challenge: (2020) African languages are underrepresented in existing natural language processing datasets, research, and tools due to lack of datasets and reproducible results.
Approach: They propose to create a dataset for named entity recognition (NER) in ten African languages.
Outcome: The results of the first large dataset for named entity recognition (NER) in ten African languages are released to inform future research on African NLP.
Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead (2025.emnlp-main)

Copied to clipboard

Challenge: African languages are often left behind in state-of-the-art natural language processing systems and large language models.
Approach: They analyze 884 research papers on NLP for African languages published over past five years . they identify key trends shaping the field and outline promising directions .
Outcome: The findings identify key trends shaping the field and outline promising directions . the authors analyze 884 research papers on NLP for African languages published over the past five years .
Towards Afrocentric NLP for African Languages: Where We Are and Where We Can Go (2022.acl-long)

Copied to clipboard

Challenge: ACL 2022 special Theme on "Language Diversity: from Low Resource to Endangered Languages" focuses on linguistic and sociopolitical challenges facing development of NLP technologies for African languages .
Approach: They propose a typological framework for linguistic and sociopolitical challenges for NLP in African languages.
Outcome: The main objective of this study is to motivate and advocate for an Afrocentric approach to technology development.
AfroBench: How Good are Large Language Models on African Languages? (2025.findings-acl)

Copied to clipboard

Challenge: Large-scale multilingual evaluations often include only a handful of African languages due to the scarcity of high-quality data and the limited discoverability of existing datasets.
Approach: They propose a multi-task benchmark to evaluate the performance of LLMs across 64 African languages, 15 tasks and 22 datasets.
Outcome: The proposed benchmark compares LLMs across 64 African languages, 15 tasks and 22 datasets.
An Open Dataset and Model for Language Identification (2023.acl-short)

Copied to clipboard

Challenge: Existing LID systems perform poorly on low-resource languages, causing 'representation washing', where the community is given a false view of the actual progress of low-source NLP.
Approach: They propose a model which achieves a macro-average F1 score of 0.93 and a false positive rate of 0.033% across 201 languages, outperforming previous work.
Outcome: The proposed model outperforms existing models and datasets on 201 languages and a false positive rate of 0.033%.
AfroMT: Pretraining Strategies and Reproducible Benchmarks for Translation of 8 African Languages (2021.emnlp-main)

Copied to clipboard

Challenge: Existing reproducible benchmarks for machine translation are limited to high-resource or well-represented languages.
Approach: They propose to use AfroMT to develop a reproducible machine translation benchmark for eight widely spoken African languages and a suite of analysis tools to take into account their unique properties.
Outcome: The proposed benchmarks show significant improvements when pretraining on 11 languages, with gains of up to 2 BLEU points over strong baselines.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations