Papers by Anuoluwapo Aremu

13 papers
ÌròyìnSpeech: A Multi-purpose Yorùbá Speech Corpus (2024.lrec-main)

Copied to clipboard

Challenge: rynSpeech corpus is a dataset that can be used for both Text-to-Speecher (TTS) and Automatic Speech Recognition (ASR) speakers of many African languages have no access to voice-enabled applications in their native languages.
Approach: They propose a dataset to collect Yorùbá speech data that can be used for both TTS and ASR tasks.
Outcome: The proposed dataset can generate a good quality model with as little as 5 hours of speech . the results are consistent with previous studies on the Yorùbá language .
Does Generative AI speak Nigerian-Pidgin?: Issues about Representativeness and Bias for Multilingualism in LLMs (2025.findings-naacl)

Copied to clipboard

Challenge: Nigeria is a multilingual country with 500+ languages.
Approach: They propose to use a pidgin and a creole to analyze the pidgins of Nigeria . they also use machine translation to analyze their results .
Outcome: The results show that the two pidgins do not represent each other and are hard to teach . the results show the pidgin varieties are underrepresented in Generative AI .
MasakhaNER: Named Entity Recognition for African Languages (2021.tacl-1)

Copied to clipboard

Challenge: (2020) African languages are underrepresented in existing natural language processing datasets, research, and tools due to lack of datasets and reproducible results.
Approach: They propose to create a dataset for named entity recognition (NER) in ten African languages.
Outcome: The results of the first large dataset for named entity recognition (NER) in ten African languages are released to inform future research on African NLP.
Multi-lingual and Multi-cultural Figurative Language Understanding (2023.findings-acl)

Copied to clipboard

Challenge: Figures permeate human communication, but are understudied in NLP.
Approach: They create a figurative language inference dataset for seven languages associated with a variety of cultures, using cultural and regional concepts for figurativ expressions.
Outcome: The results show that the most common figurative expressions are found in Hindi, Indonesian, Javanese, Kannada, Sundanese, Swahili and Yoruba.
AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African Languages (2024.naacl-long)

Copied to clipboard

Challenge: Recent advances in machine translation (MT) have focused on scaling multilingual machine translation models and evaluation data to hundreds of languages, including multiple under-resourced languages.
Approach: They propose to use n-gram matching metrics to measure progress in multilingual machine translation to 13 typologically diverse African languages to create high-quality human evaluation data with simplified MQM guidelines.
Outcome: The proposed metrics have a higher correlation with human judgments than n-gram matching metrics such as BLEU and METEOR.
The Esethu Framework: Reimagining Sustainable Dataset Governance and Curation for Low-Resource Languages (2025.acl-long)

Copied to clipboard

Challenge: Esethu Framework is a community-centric data license that empowers local communities and ensures equitable benefit-sharing from their linguistic resource.
Approach: They propose a community-centric data license to empower local communities and ensure equitable benefit-sharing from their linguistic resource.
Outcome: The proposed dataset contains read speech from native isiXhosa speakers enriched with demographic and linguistic metadata.
NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Sentiment analysis is one of the most widely studied applications in NLP, but most work focuses on languages with large amounts of data.
Approach: They propose a large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria.
Outcome: The proposed dataset includes 30,000 tweets and a significant fraction of code-mixed tweets.
Mitigating Translationese in Low-resource Languages: The Storyboard Approach (2024.lrec-main)

Copied to clipboard

Challenge: Low-resource languages often face challenges in acquiring high-quality language data due to the reliance on translation-based methods, which introduce the translationese effect.
Approach: They propose a method that uses storyboards to elicit more fluent and natural sentences from native speakers without direct exposure to the source text.
Outcome: The proposed method compared with traditional translation-based methods in terms of accuracy and fluency.
Voices Unheard: NLP Resources and Models for Yorùbá Regional Dialects (2024.emnlp-main)

Copied to clipboard

Challenge: Recent efforts to develop NLP tools for low-resource languages focus on their standard dialects.
Approach: They propose a high-quality parallel text and speech corpus for Yoruba . they use native speakers to collect data from four regional yoruba dialects .
Outcome: The proposed dataset shows that dialect-adaptive finetuning can narrow performance disparities . the dataset will be released publicly under an open license .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations