Papers by Odunayo Ogundepo

10 papers
Better Quality Pre-training Data and T5 Models for African Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Existing web crawls have demonstrated quality issues for low-resource languages . Existing pretraining corpora have numerous quality issues .
Approach: They propose to audit existing pretraining corpora to understand and rectify quality issues . they pretrain a new T5-based model and evaluate its performance on multiple tasks .
Outcome: The proposed model outperforms existing pretrained models on four NLP tasks.
AfriCLIRMatrix: Enabling Cross-Lingual Information Retrieval for African Languages (2022.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for cross-lingual information retrieval are limited in many languages, especially those spoken in Africa.
Approach: They propose to build a test collection for cross-lingual information retrieval in 15 diverse African languages.
Outcome: AfriCLIRMatrix contains 6 million queries in English and 23 million relevance judgments automatically mined from Wikipedia inter-language links, covering many more African languages than any existing information retrieval test collection.
“Knowing When You Don’t Know”: A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation (2024.findings-emnlp)

Copied to clipboard

Challenge: Prior work on RAG grounds Large Language Models to reduce factual hallucinations lacks a comprehensive evaluation of different language families.
Approach: They propose a human-annotated dataset for evaluating LLM robustness in RAG . they find that most models struggle to balance the two capacities .
Outcome: The proposed dataset includes both a non-relevant and a relevant subset.
MIRACL: A Multilingual Retrieval Dataset Covering 18 Diverse Languages (2023.tacl-1)

Copied to clipboard

Challenge: MIRACL is a multilingual dataset for ad hoc retrieval across 18 languages that collectively encompass over three billion native speakers around the world.
Approach: They have gathered over 726k high-quality relevance judgments for 78k queries over Wikipedia in these languages, where all annotations have been performed by native speakers hired by their team.
Outcome: MIRACL covers languages that are typologically close as well as distant from 10 language families and 13 sub-families, associated with varying amounts of publicly available resources.
Spacerini: Plug-and-play Search Engines with Pyserini and Hugging Face (2023.emnlp-demo)

Copied to clipboard

Challenge: a toolkit for reproducible information retrieval research is available for free.
Approach: They present a tool that integrates Pyserini and Hugging Face to enable the seamless construction and deployment of interactive search engines.
Outcome: The proposed tool makes state-of-the-art retrieval models more accessible to non-IR practitioners while minimizing deployment effort.
Evaluating Embedding APIs for Information Retrieval (2023.acl-industry)

Copied to clipboard

Challenge: a growing number of language models are limiting their access to the community . we evaluate existing APIs for domain generalization and multilingual retrieval .
Approach: They evaluate semantic embedding APIs in retrieval scenarios to assess their capabilities . they use BEIR and MIRACL to re-rank BM25 results using the APIs .
Outcome: The proposed model is based on semantic embedding APIs that build vector representations of a given text.
GAIA Search: Hugging Face and Pyserini Interoperability for NLP Training Data Exploration (2023.acl-demo)

Copied to clipboard

Challenge: Using the mature and well-tested methods from the domain of Information Retrieval (IR) we propose to integrate Pyserini with Hugging Face to provide qualitative analysis tools for NLP researchers.
Approach: They propose to integrate Pyserini with Hugging Face to provide qualitative analysis tools for NLP researchers.
Outcome: The proposed tools can be integrated with the Hugging Face ecosystem of open-source AI libraries and artifacts.
AfroBench: How Good are Large Language Models on African Languages? (2025.findings-acl)

Copied to clipboard

Challenge: Large-scale multilingual evaluations often include only a handful of African languages due to the scarcity of high-quality data and the limited discoverability of existing datasets.
Approach: They propose a multi-task benchmark to evaluate the performance of LLMs across 64 African languages, 15 tasks and 22 datasets.
Outcome: The proposed benchmark compares LLMs across 64 African languages, 15 tasks and 22 datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations