Papers by Ife Adebara
Fumbling in Babel: An Investigation into ChatGPT’s Language Identification Ability (2024.findings-naacl)
Copied to clipboard
| Challenge: | ChatGPT is a powerful NLP tool but its language identification abilities are unclear. |
| Approach: | They compile a benchmark comprising 670 languages representing 23 language families spoken in five continents and compare their language identification abilities to ChatGPT's (both GPT-3.5 and GPT-4) performance. |
| Outcome: | The proposed model performs poorly on African languages, while GPT-3.5 and GPT-4 perform poorly on English, Afrikaans, Arabic, Indonesian, Italian, Mandarin Chinese, and several more. |
Linguistically-Motivated Yorùbá-English Machine Translation (2022.coling-1)
Copied to clipboard
| Challenge: | Several phenomena where asymmetry arises have been identified as challenging problems for machine translation. |
| Approach: | They perform a fine-grained analysis of how an SMT system compares with two NMT systems when translating bare nouns into English. |
| Outcome: | The proposed model outperforms the SMT and BiLSTM models for 4 categories and the BiLST outperformed the SLT models for 3 categories. |
Cheetah: Natural Language Generation for 517 African Languages (2024.acl-long)
Copied to clipboard
| Challenge: | Low-resource African languages pose unique challenges for natural language processing (NLG) We demonstrate the effectiveness of Cheetah through comprehensive evaluations across six generation downstream tasks. |
| Approach: | They develop a multilingual NLG language model for African languages called Cheetah . they demonstrate that Cheethah outperforms other models in six tasks . |
| Outcome: | The proposed model outperforms other models in five of six generation tasks. |
Towards Afrocentric NLP for African Languages: Where We Are and Where We Can Go (2022.acl-long)
Copied to clipboard
| Challenge: | ACL 2022 special Theme on "Language Diversity: from Low Resource to Endangered Languages" focuses on linguistic and sociopolitical challenges facing development of NLP technologies for African languages . |
| Approach: | They propose a typological framework for linguistic and sociopolitical challenges for NLP in African languages. |
| Outcome: | The main objective of this study is to motivate and advocate for an Afrocentric approach to technology development. |
SERENGETI: Massively Multilingual Language Models for Africa (2023.findings-acl)
Copied to clipboard
| Challenge: | Pretrained models acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning. |
| Approach: | They develop a set of massively multilingual language models that covers 517 African languages and language varieties. |
| Outcome: | The proposed models outperform 4 models that cover 4-23 African languages on eight natural language understanding tasks, achieving 82.27 average F_1. |
Toucan: Many-to-Many Translation for 150 African Language Pairs (2024.findings-acl)
Copied to clipboard
| Challenge: | We introduce two language models with 1.2 billion and 3.7 billion parameters to improve Machine Translation (MT) for low-resource languages. |
| Approach: | They propose a set of tools to improve Machine Translation (MT) for low-resource languages with a focus on African languages. |
| Outcome: | The proposed model outperforms existing models on MT for African languages and improves translation evaluation metrics for 1K languages including African languages. |
Interplay of Machine Translation, Diacritics, and Diacritization (2024.naacl-long)
Copied to clipboard
| Challenge: | MT and diacritization influence performance in a multi-task learning setting, but keeping diacritics is harmful for some languages. |
| Approach: | They propose two classes of metrics to measure the complexity of a diacritical system and propose to use them to compare performance. |
| Outcome: | The proposed metrics correlate positively with the performance of the models. |
Where Are We? Evaluating LLM Performance on African Languages (2025.acl-long)
Copied to clipboard
Ife Adebara, Hawau Olamide Toyin, Nahom Tesfu Ghebremichael, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed
| Challenge: | African languages are underrepresented in NLP due to policies that favor foreign languages and create data inequities. |
| Approach: | They integrate theoretical insights on Africa’s language landscape with an empirical evaluation using Sahara datasets. |
| Outcome: | The proposed model improves on a benchmark curated from large-scale, publicly accessible datasets capturing the continent's linguistic diversity. |
AfroLID: A Neural Language Identification Tool for African Languages (2022.emnlp-main)
Copied to clipboard
| Challenge: | AfroLID is a neural LID toolkit for 517 African languages and varieties. |
| Approach: | They propose to exploit a multi-domain web dataset manually curated from across 14 language families utilizing five orthographic systems to exploit AfroLID. |
| Outcome: | The proposed tool outperforms existing tools on the acutely under-served Twitter domain. |