Papers by Kelechi Ogueji

7 papers
MasakhaNER: Named Entity Recognition for African Languages (2021.tacl-1)

Copied to clipboard

Challenge: (2020) African languages are underrepresented in existing natural language processing datasets, research, and tools due to lack of datasets and reproducible results.
Approach: They propose to create a dataset for named entity recognition (NER) in ten African languages.
Outcome: The results of the first large dataset for named entity recognition (NER) in ten African languages are released to inform future research on African NLP.
What a Creole Wants, What a Creole Needs (2022.lrec-1)

Copied to clipboard

Challenge: Recent efforts to improve the quality of high-resource languages focus on translating existing datasets into other languages, but this approach ignores that different language communities have different needs.
Approach: They examine how things needed from language technology can change dramatically from one language to another.
Outcome: The proposed method ignores that different language communities have different needs.
Intriguing Properties of Compression on Multilingual Models (2022.emnlp-main)

Copied to clipboard

Challenge: Multilingual models are dependent on scaling to generalize to a growing number of languages . compression techniques can have disparate effects on model performance for low-resource languages if used sparsely .
Approach: They propose to characterize the impact of sparsifying multilingual pre-trained language models during fine-tuning.
Outcome: The proposed framework characterizes the impact of sparsifying multilingual pre-trained language models during fine-tuning.
Enhancing Alignment using Curriculum Learning & Ranked Preferences (2024.findings-emnlp)

Copied to clipboard

Challenge: Direct Preference Optimization (DPO) is an effective technique that leverages pairwise preference data to align LLMs to human preferences.
Approach: They propose to use pairwise preference data to create multiple preference pairs for a given prompt.
Outcome: The proposed method outperforms standard DPO on MTbench, Vicuna bench, and WizardLM with a score of 7.43 on the test sets.
AfroBench: How Good are Large Language Models on African Languages? (2025.findings-acl)

Copied to clipboard

Challenge: Large-scale multilingual evaluations often include only a handful of African languages due to the scarcity of high-quality data and the limited discoverability of existing datasets.
Approach: They propose a multi-task benchmark to evaluate the performance of LLMs across 64 African languages, 15 tasks and 22 datasets.
Outcome: The proposed benchmark compares LLMs across 64 African languages, 15 tasks and 22 datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations