Papers by Abraham Owodunni

4 papers
Better Quality Pre-training Data and T5 Models for African Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Existing web crawls have demonstrated quality issues for low-resource languages . Existing pretraining corpora have numerous quality issues .
Approach: They propose to audit existing pretraining corpora to understand and rectify quality issues . they pretrain a new T5-based model and evaluate its performance on multiple tasks .
Outcome: The proposed model outperforms existing pretrained models on four NLP tasks.
Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks (2022.emnlp-main)

Copied to clipboard

Challenge: In total, the initial release of the Bloom Library datasets covers 363 languages across 32 language families.
Approach: They present a set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/recognition.
Outcome: The Bloom Library datasets cover 363 languages across 32 language families.
AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents (2024.findings-eacl)

Copied to clipboard

Challenge: AccentFold uses spatial relationships to improve speech recognition for accented speech . existing methods for accent recognition have been limited due to data scarcity and budget constraints .
Approach: They propose a method that exploits spatial relationships between learned accent embeddings to improve downstream automatic speech recognition.
Outcome: The proposed method outperforms baseline methods in accented speech training.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations