Challenge: Illiteracy is a predictor of many negative social and personal outcomes in underresourced countries, where few books exist that are suitable for children to learn to read from.
Approach: They propose to use generative AI to create culturally-engaging materials for learning in mali's vehicular language Bambara by multiplying the content by 10 times . authors propose to apply bias-aware tools to reduce illiteracy and improve learning outcomes through native language education.
Outcome: The proposed toolchain and workflow can be adapted to address low literacy in mali using generative AI.

Similar Papers

Learnings from Technological Interventions in a Low Resource Language: A Case-Study on Gondi (2020.lrec-1)

Copied to clipboard

Challenge: 40% of all the languages in the world face the danger of extinction in the near future . when a language dies out, future generations lose a vital part of the culture that is necessary to completely understand it.
Approach: They propose to use 4 technology-driven methods of data collection to collect data on Gondi, a low-resource vulnerable language spoken by 2.3 million tribal people in south and central India.
Outcome: The proposed methods collected 12,000 translated words and/or sentences and identified more than 650 community members whose help can be solicited for future translation efforts.
Building Representative Corpora from Illiterate Communities: A Reviewof Challenges and Mitigation Strategies for Developing Countries (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods for collecting data from high-income countries (HICs) make implicit assumptions about literacy and internet access, but in low-income and sub-Saharan Africa (SSA) such assumptions may not hold for LICs where the bulk of the population lives.
Approach: They propose a set of practical mitigation strategies to address the under-representation of illiterate communities in NLP corpora.
Outcome: The proposed methods address the under-representation of illiterate communities in NLP corpora and propose mitigation strategies to help future work.
Investigating Meta-Learning Algorithms for Low-Resource Natural Language Understanding Tasks (D19-1)

Copied to clipboard

Challenge: Existing methods to learn general representations of text can achieve sub-optimal performance in low-resource scenarios.
Approach: They propose to use language model pre-training and multi-task learning to learn robust representations but these methods can achieve sub-optimal performance in low-resource scenarios.
Outcome: The proposed model outperforms strong baselines on the GLUE benchmark and can be adapted to new tasks efficiently and effectively.
GameQA: Gamified Mobile App Platform for Building Multiple-Domain Question-Answering Datasets (2023.eacl-demo)

Copied to clipboard

Challenge: a common problem with question-answering datasets is that they require annotators to source answers from the internet . a crowd-sourcing platform is available for low-resource languages, but it is limited in terms of information available.
Approach: They propose a crowd-sourcing platform to gather multiple-domain QA data for low-resource languages.
Outcome: The proposed platform rivals large QA datasets for high-resource languages in size and answerability.
Counterspeech Generation using Small Language Models (2026.acl-srw)

Copied to clipboard

Challenge: Social media use is growing annually with about 68.5% of the global population active on these platforms as of July 2025.
Approach: They evaluate SLMs ranging from 100 million to 3 billion parameters using simple prompting strategies as well as fine-tuning, combining automatic and robust human evaluations.
Outcome: The proposed models generate relevant, coherent, and high-quality counterspeech, suggesting their suitability for efficient and responsible deployments.
Scaling Cultural Resources for Improving Generative Models (2026.findings-eacl)

Copied to clipboard

Challenge: generative models have been known to have reduced performance in different global cultural contexts and languages.
Approach: They construct a pipeline to collect and contribute culturally salient, multilingual data . they argue such data can assess the state of the global applicability of generative AI models .
Outcome: The proposed pipeline can assess the state of the global applicability of our models and improve upon cross-cultural gaps.
Multilingual Dependency Parsing for Low-Resource Languages: Case Studies on North Saami and Komi-Zyrian (L18-1)

Copied to clipboard

Challenge: Developing systems for low-resource languages is a crucial issue for Natural Language Processing (NLP).
Approach: They propose a method for parsing low-resource languages with very small training corpora using multilingual word embeddings and annotated corporata of larger languages.
Outcome: The proposed method improves dependency parsing for low-resource languages with very small training corpora compared to previous work . it also explores whether contemporary contact languages or genetically related languages would be the most fruitful starting point for multilingual parsers.
SAFARI: A Community-Engaged Approach and Dataset of Stereotype Resources in the Sub-Saharan African Context (2026.eacl-short)

Copied to clipboard

Challenge: Existing data collection approaches to generative AI are inadequate to assess its safety and utility.
Approach: They propose a multilingual stereotype resource that uses socioculturally-situated, community-engaged methods to assess the region’s linguistic diversity and traditional orality.
Outcome: The proposed method covers four sub-Saharan African countries that are severely underrepresented in NLP resources: Ghana, Kenya, Nigeria, and South Africa.
Language-Agnostic Meta-Learning for Low-Resource Text-to-Speech with Articulatory Features (2022.acl-long)

Copied to clipboard

Challenge: Recent advances in text-to-speech systems allow for speech synthesis with unprecedented quality and controllability.
Approach: They use embeddings derived from articulatory vectors rather than phoneme identities to learn phoneme representations that hold across languages.
Outcome: The proposed models fine-tuned on 30 minutes of data in a previously unseen language with language agnostic meta learning.
Systematic Investigation of Strategies Tailored for Low-Resource Settings for Low-Resource Dependency Parsing (2023.eacl-main)

Copied to clipboard

Challenge: Several strategies have been proposed to enhance performance in low-resource scenarios.
Approach: They propose to use 5 low-resource strategies for dependency parsing for multiple languages . they use ensembled approach on 7 UD low-rsource languages based on their results .
Outcome: The proposed approach improves on a low-resource language Sanskrit.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations