Papers by Ife Adebara

9 papers
Fumbling in Babel: An Investigation into ChatGPT’s Language Identification Ability (2024.findings-naacl)

Copied to clipboard

Challenge: ChatGPT is a powerful NLP tool but its language identification abilities are unclear.
Approach: They compile a benchmark comprising 670 languages representing 23 language families spoken in five continents and compare their language identification abilities to ChatGPT's (both GPT-3.5 and GPT-4) performance.
Outcome: The proposed model performs poorly on African languages, while GPT-3.5 and GPT-4 perform poorly on English, Afrikaans, Arabic, Indonesian, Italian, Mandarin Chinese, and several more.
Linguistically-Motivated Yorùbá-English Machine Translation (2022.coling-1)

Copied to clipboard

Challenge: Several phenomena where asymmetry arises have been identified as challenging problems for machine translation.
Approach: They perform a fine-grained analysis of how an SMT system compares with two NMT systems when translating bare nouns into English.
Outcome: The proposed model outperforms the SMT and BiLSTM models for 4 categories and the BiLST outperformed the SLT models for 3 categories.
Cheetah: Natural Language Generation for 517 African Languages (2024.acl-long)

Copied to clipboard

Challenge: Low-resource African languages pose unique challenges for natural language processing (NLG) We demonstrate the effectiveness of Cheetah through comprehensive evaluations across six generation downstream tasks.
Approach: They develop a multilingual NLG language model for African languages called Cheetah . they demonstrate that Cheethah outperforms other models in six tasks .
Outcome: The proposed model outperforms other models in five of six generation tasks.
Towards Afrocentric NLP for African Languages: Where We Are and Where We Can Go (2022.acl-long)

Copied to clipboard

Challenge: ACL 2022 special Theme on "Language Diversity: from Low Resource to Endangered Languages" focuses on linguistic and sociopolitical challenges facing development of NLP technologies for African languages .
Approach: They propose a typological framework for linguistic and sociopolitical challenges for NLP in African languages.
Outcome: The main objective of this study is to motivate and advocate for an Afrocentric approach to technology development.
SERENGETI: Massively Multilingual Language Models for Africa (2023.findings-acl)

Copied to clipboard

Challenge: Pretrained models acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning.
Approach: They develop a set of massively multilingual language models that covers 517 African languages and language varieties.
Outcome: The proposed models outperform 4 models that cover 4-23 African languages on eight natural language understanding tasks, achieving 82.27 average F_1.
Toucan: Many-to-Many Translation for 150 African Language Pairs (2024.findings-acl)

Copied to clipboard

Challenge: We introduce two language models with 1.2 billion and 3.7 billion parameters to improve Machine Translation (MT) for low-resource languages.
Approach: They propose a set of tools to improve Machine Translation (MT) for low-resource languages with a focus on African languages.
Outcome: The proposed model outperforms existing models on MT for African languages and improves translation evaluation metrics for 1K languages including African languages.
Interplay of Machine Translation, Diacritics, and Diacritization (2024.naacl-long)

Copied to clipboard

Challenge: MT and diacritization influence performance in a multi-task learning setting, but keeping diacritics is harmful for some languages.
Approach: They propose two classes of metrics to measure the complexity of a diacritical system and propose to use them to compare performance.
Outcome: The proposed metrics correlate positively with the performance of the models.
Where Are We? Evaluating LLM Performance on African Languages (2025.acl-long)

Copied to clipboard

Challenge: African languages are underrepresented in NLP due to policies that favor foreign languages and create data inequities.
Approach: They integrate theoretical insights on Africa’s language landscape with an empirical evaluation using Sahara datasets.
Outcome: The proposed model improves on a benchmark curated from large-scale, publicly accessible datasets capturing the continent's linguistic diversity.
AfroLID: A Neural Language Identification Tool for African Languages (2022.emnlp-main)

Copied to clipboard

Challenge: AfroLID is a neural LID toolkit for 517 African languages and varieties.
Approach: They propose to exploit a multi-domain web dataset manually curated from across 14 language families utilizing five orthographic systems to exploit AfroLID.
Outcome: The proposed tool outperforms existing tools on the acutely under-served Twitter domain.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations