Papers by Muhammad Abdul-Mageed
Copied to clipboard
| Challenge: | Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. |
| Approach: | They propose to construct a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding. |
| Outcome: | The proposed dataset covers ten culturally significant domains covering all Arab countries and includes two evaluation benchmarks (PEARL and PEARL-LITE) and a specialized subset (PearL-X). |
Copied to clipboard
| Challenge: | Large-scale multilingual ASR has substantially improved recognition for high-resource languages. |
| Approach: | They propose a proxy-guided -best selection paradigm that conditions inference on external side information without parameter updates. |
| Outcome: | The proposed model reduces WER by 15.6% relative and recovers a fraction of oracle n-best gains on the common voice MSA testbed. |
Copied to clipboard
| Challenge: | MLLMs have proven effective in a wide range of tasks that require complex reasoning and linguistic comprehension, but they are limited to English-based settings. |
| Approach: | They propose a family of Arabic multimodal large language models with strong vision and language capabilities. |
| Outcome: | The proposed models show strong performance on visual reasoning tasks and language capabilities. |
Copied to clipboard
| Challenge: | Neural architecture search (NAS) uses weight-sharing supernets to generate diverse subnetworks without retraining. |
| Approach: | They propose a weight-sharing supernet that leverages mixture-of-experts to enhance supernet model expressiveness with minimal training overhead. |
| Outcome: | The proposed method achieves state-of-the-art (SoTA) performance in NAS for fast machine translation models, surpassing NAS-BERT and AutoDistil across various model sizes. |
Copied to clipboard
| Challenge: | Existing approaches to fine-tune pre-trained language models for downstream tasks require labeled data. |
| Approach: | They propose to self-train pre-trained language models to improve performance on data-scarce varieties by as large as 10% F1 and 2% accuracy. |
| Outcome: | The proposed model improves zero-shot MSA-to-DA transfer by as large as 10% F1 (NER) and 2% accuracy (POS tagging). |
Copied to clipboard
| Challenge: | a large reasoning model (LRM) training on large amounts of reasoning data is computationally expensive. |
| Approach: | They propose a method to quantify computation-quality tradeoffs as a function of sequence length. |
| Outcome: | The proposed method reduces training time, memory and FLOPs by 50% on long training sequences while retaining the full-sequence performance. |
Copied to clipboard
| Challenge: | Our study examines ChatGPT’s performance on Arabic languages and dialectal varieties. |
| Approach: | They conduct a large-scale automated and human evaluation of ChatGPT, encompassing 44 distinct language understanding and generation tasks on over 60 different datasets. |
| Outcome: | The proposed model outperforms smaller models on Arabic dialects compared to GPT-4's Modern Standard Arabic and Dialectal Arabic (DA) |
Copied to clipboard
| Challenge: | Existing MoE designs do not consider computational constraints (e.g., FLOPs, latency) Existing works in MoE consider homogeneous design where the same number of experts of the same size are placed uniformly throughout the network. |
| Approach: | They propose a framework for designing heterogeneous MoEs under computational constraints. |
| Outcome: | The proposed framework achieves 4x inference speedup and FLOPs reduction over manual models and within 1 BLEU point of MoE SwitchTransformer over benchmark datasets for NMT. |
Copied to clipboard
| Challenge: | ChatGPT is a powerful NLP tool but its language identification abilities are unclear. |
| Approach: | They compile a benchmark comprising 670 languages representing 23 language families spoken in five continents and compare their language identification abilities to ChatGPT's (both GPT-3.5 and GPT-4) performance. |
| Outcome: | The proposed model performs poorly on African languages, while GPT-3.5 and GPT-4 perform poorly on English, Afrikaans, Arabic, Indonesian, Italian, Mandarin Chinese, and several more. |
Copied to clipboard
| Challenge: | Despite efforts to evaluate Arabic NLU, no public benchmark of diverse nature exists . a benchmark targeting Arabic needs to take into account that Arabic is not a single language but a collection of languages and language varieties. |
| Approach: | They propose a publicly available benchmark for Arabic language understanding evaluation dubbed ORCA . it covers diverse Arabic varieties and a wide range of Arabic understanding tasks . |
| Outcome: | The proposed benchmark covers Arabic and multilingual models across seven NLU task clusters. |
Copied to clipboard
| Challenge: | Several phenomena where asymmetry arises have been identified as challenging problems for machine translation. |
| Approach: | They perform a fine-grained analysis of how an SMT system compares with two NMT systems when translating bare nouns into English. |
| Outcome: | The proposed model outperforms the SMT and BiLSTM models for 4 categories and the BiLST outperformed the SLT models for 3 categories. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated impressive performance on a wide range of natural language processing tasks. |
| Approach: | They propose an unsupervised approach to mine in-context examples for machine translation (MT) they use word-level mining to acquire word translations that are then used to perform sentence-level mines . |
| Outcome: | The proposed approach outperforms state-of-the-art methods on 288 directions on 287 languages and is based on word-level mining and sentence-level extraction. |
Copied to clipboard
| Challenge: | In this paper, we introduce a family of embedding models addressing both small-scale and large-scale use cases. |
| Approach: | They propose to use ArabicMTEB to evaluate Arabic text embedding models . they propose to build a benchmark suite that assesses cross-lingual, multi-dialectal, multidomain, and multi-cultural Arabic text embedded models. |
| Outcome: | The proposed models outperform Multilingual-E5-large and Swan-Large in most Arabic tasks while remaining dialectally and culturally aware. |
Copied to clipboard
| Challenge: | a global pandemic of coronavirus disease 2019 has impacted millions of people . a human annotation study reveals the utility of our models on a subset of Mega-COV . |
| Approach: | They develop powerful models to analyze tweets related to the pandemic . they use a multilingual Twitter dataset with geo-location information . |
| Outcome: | The proposed model can identify whether a tweet is related to the pandemic and detect misinformation about it. |
Copied to clipboard
| Challenge: | Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than modern standard Arabic (MSA). |
| Approach: | They propose a large-scale, community-driven, human-translated dataset to bridge this gap . Alexandria covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . |
| Outcome: | The Alexandria dataset covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . Alexandria is a training resource and a rigorous benchmark for evaluating MT and LLMs based on the Alexandria dataset . |
Copied to clipboard
| Challenge: | Low-resource African languages pose unique challenges for natural language processing (NLG) We demonstrate the effectiveness of Cheetah through comprehensive evaluations across six generation downstream tasks. |
| Approach: | They develop a multilingual NLG language model for African languages called Cheetah . they demonstrate that Cheethah outperforms other models in six tasks . |
| Outcome: | The proposed model outperforms other models in five of six generation tasks. |
Copied to clipboard
| Challenge: | linguistic diversity in Africa is underrepresented in speech technologies, creating barriers to digital inclusion. |
| Approach: | They propose a benchmarking framework to map the continent's linguistic diversity and map its impact on downstream African speech tasks. |
| Outcome: | The proposed model achieves state-of-the-art across multiple African languages and speech tasks. |
Copied to clipboard
| Challenge: | Dialectal Arabic datasets embody a range of domain, dialect, and quality. |
| Approach: | They propose a framework for automatic speech recognition in dialectal Arabic to address the limited data availability encountered in dialects. |
| Outcome: | The proposed framework provides access to 31 datasets covering 14 dialects to better address the limited data availability encountered in dialectal Arabic speech processing. |
Copied to clipboard
| Challenge: | Recent advances in instruction fine-tuning and alignment methods have enhanced the adaptability of large language models to user preferences. |
| Approach: | They propose a benchmark to assess LLMs’ capacity to comprehend and interpret Arabic proverbs. |
| Outcome: | The proposed model can generate accurate translations, but struggle to produce culturally nuanced and contextually relevant explanations. |
Copied to clipboard
| Challenge: | grammatical gender significantly influences image generation in text-to-image models . masculine grammatikal markers increase male representation to 73% on average . feminine grammatological markers increase female representation to 38% . |
| Approach: | They propose a cross-linguistic benchmark examining words where grammatical gender contradicts stereotypical gender associations. |
| Outcome: | The proposed benchmark examines words where grammatical gender contradicts stereotypical gender associations. |
Copied to clipboard
| Challenge: | Existing instruction tuned large language models (LLMs) struggle to understand cross-lingual sociopragmatic meaning (SM) lack of comprehensive investigation into their ability to understand SM is partly due to SM not being adequately represented in any of the existing benchmarks. |
| Approach: | They evaluate the performance of instruction tuned large language models (LLMs) on a multilingual benchmark specifically designed for SM understanding. |
| Outcome: | The proposed benchmark outperforms instruction tuned large language models on a wide range of tasks but falls behind task-specific finetuned models. |
Copied to clipboard
| Challenge: | Existing models for multilingual automatic speech recognition (ASR) are computationallyintensive and lack proper comprehensive evaluations. |
| Approach: | They propose to distill knowledge from large teacher models into smaller student variants that are more efficient. |
| Outcome: | The proposed model outperforms existing models on standard benchmarks and dialectal data. |
Copied to clipboard
| Challenge: | Existing studies of Arabic dialects have focused on blogs and comments on online news sites, but data on other dialects are costly and limited. |
| Approach: | They present a dataset of > 1/4 billion tweets representing a wide range of Arabic dialects. |
| Outcome: | The dataset represents 29 major Arab cities from 10 Arab countries with varying dialects. |
Copied to clipboard
| Challenge: | DetoxLLM is a comprehensive end-to-end detoxification framework for toxic language. |
| Approach: | They propose a comprehensive end-to-end detoxification framework that tackles toxic language across platforms. |
| Outcome: | The proposed detoxification framework outperforms the SoTA model on human-annotated parallel corpus and offers explanation to promote transparency and trustworthiness. |
Copied to clipboard
| Challenge: | Large language models with instruction tuning are resource-intensive . a recent study suggests that the performance of LLMs scales proportionally with the size of the model. |
| Approach: | They propose to distill knowledge from instruction-tuned LLMs into much smaller ones . they develop a large set of 2.58M instructions based on existing and newly-generated instructions . |
| Outcome: | The proposed models are comparable to strong baselines while being much smaller in size. |
Copied to clipboard
| Challenge: | ACL 2022 special Theme on "Language Diversity: from Low Resource to Endangered Languages" focuses on linguistic and sociopolitical challenges facing development of NLP technologies for African languages . |
| Approach: | They propose a typological framework for linguistic and sociopolitical challenges for NLP in African languages. |
| Outcome: | The main objective of this study is to motivate and advocate for an Afrocentric approach to technology development. |
Copied to clipboard
| Challenge: | Current research directions rely on synthetic data generated by translating English corpora, which often fails to represent the cultural heritage and values of local communities. |
| Approach: | They propose a method to create and retrieve pre-training data tailored to a specific community . they use Egyptian and Moroccan dialects as testbeds to test their understanding . |
| Outcome: | The proposed method outperforms existing Arabic-aware LLMs and performs on par with larger models. |
Copied to clipboard
| Challenge: | a year-long community-driven project covering all 22 Arab countries evaluates the cultural and dialectal capabilities of large language models. |
| Approach: | They propose a project to evaluate the cultural and dialectal capabilities of large language models. |
| Outcome: | The project evaluates the cultural and dialectal capabilities of several frontier LLMs. |
Copied to clipboard
| Challenge: | Pretrained models acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning. |
| Approach: | They develop a set of massively multilingual language models that covers 517 African languages and language varieties. |
| Outcome: | The proposed models outperform 4 models that cover 4-23 African languages on eight natural language understanding tasks, achieving 82.27 average F_1. |
Copied to clipboard
| Challenge: | Recent advances in generative AI have transformed the landscape of writing assistance, especially through the development of Large Language Models (LLMs). |
| Approach: | They propose to use a dataset to evaluate leading LLMs to improve their writing assistance tools in Arabic. |
| Outcome: | The proposed dataset highlights the strengths and limitations of leading LLMs, including GPT-**4**, GPT**4o**, Cohere Command R+, and Gemini **1.5** Pro. |
Copied to clipboard
| Challenge: | Existing benchmarks for Arabic are limited, but they can be used to measure performance of different languages. |
| Approach: | They propose a benchmark for Arabic that addresses the need for a framework dedicated to Arabic languages and varieties. |
| Outcome: | The proposed benchmark covers 13 different tasks in Arabic and spans 50 test splits. |
Copied to clipboard
| Challenge: | We introduce two language models with 1.2 billion and 3.7 billion parameters to improve Machine Translation (MT) for low-resource languages. |
| Approach: | They propose a set of tools to improve Machine Translation (MT) for low-resource languages with a focus on African languages. |
| Outcome: | The proposed model outperforms existing models on MT for African languages and improves translation evaluation metrics for 1K languages including African languages. |
Copied to clipboard
| Challenge: | MT and diacritization influence performance in a multi-task learning setting, but keeping diacritics is harmful for some languages. |
| Approach: | They propose two classes of metrics to measure the complexity of a diacritical system and propose to use them to compare performance. |
| Outcome: | The proposed metrics correlate positively with the performance of the models. |
Copied to clipboard
| Challenge: | Existing MT evaluation frameworks fail to capture dialect- and culture-specific errors in diglossic languages. |
| Approach: | They propose a hierarchical error taxonomy for diagnosing MT errors through six linguistic levels: sociolinguistics, pragmatics, semantics, morphosyntax, orthography, and graphetics. |
| Outcome: | The proposed framework produces 6,113 labeled error spans across 3,495 unique erroneous sentences . it is language-agnostic and can be easily applied to or adapted for other languages. |
Copied to clipboard
| Challenge: | FinTral is a suite of state-of-the-art multimodal large language models (LLMs) built upon the Mistral-7b model and tailored for financial analysis. |
| Approach: | They introduce FinTral, a suite of state-of-the-art multimodal large language models built upon the Mistral-7b model and tailored for financial analysis. |
| Outcome: | The proposed model outperforms ChatGPT-3.5 and GPT-4 in five out of nine tasks and surpasses GPT-4.5 in five of nine task evaluations. |
Copied to clipboard
| Challenge: | African languages are underrepresented in NLP due to policies that favor foreign languages and create data inequities. |
| Approach: | They integrate theoretical insights on Africa’s language landscape with an empirical evaluation using Sahara datasets. |
| Outcome: | The proposed model improves on a benchmark curated from large-scale, publicly accessible datasets capturing the continent's linguistic diversity. |
Copied to clipboard
| Challenge: | Recent progress in representation and contrastive learning in NLP has not considered the class of sociopragmatic meaning (i.e., meaning in interaction within different language communities). |
| Approach: | They propose a framework for learning task-agnostic representations transferable to a wide range of sociopragmatic tasks. |
| Outcome: | The proposed framework outperforms other contrastive learning frameworks for both in-domain and out-of-domain data, across both the general and few-shot settings. |
Copied to clipboard
| Challenge: | Text Style Transfer (TST) aims to alter the style of text while preserving its core content. |
| Approach: | They propose a framework that leverages large language models alongside chain-of-thought prompting to facilitate TST. |
| Outcome: | The proposed framework surpasses supervised fine-tuning and knowledge distillation methods in low-resource settings. |
Copied to clipboard
| Challenge: | Existing models that convert text-based language problems into text-to-text format are not suitable for multilingual tasks. |
| Approach: | They propose a unified Transformer framework that converts all language problems into a text-to-text format. |
| Outcome: | The proposed model performs better on all ARGEN tasks than existing models with 49 less data. |
Copied to clipboard
| Challenge: | despite recent advances in speech processing, the majority of world languages and dialects remain uncovered. |
| Approach: | They propose to collect and transcribe a new Arabic dataset for eight dialects . they also develop strong baselines exploiting the new dataset . |
| Outcome: | The proposed dataset covers eight Arabic dialects, including Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni. |
Copied to clipboard
| Challenge: | Pre-trained language models (LMs) are expensive and limited in inference time . a new benchmark for multi-dialectal Arabic language understanding evaluation is developed . |
| Approach: | They introduce two powerful deep bidirectional transformer-based models, ARBERT and MARBERT . they also introduce ARLUE, a new benchmark for multi-dialectal Arabic language understanding evaluation . |
| Outcome: | The proposed models outperform monolingual models with larger vocabulary and larger datasets in Arabic language understanding evaluation. |
Copied to clipboard
| Challenge: | Current fake news detectors that exploit stylometric signals from the text are insufficient for distinguishing manipulated text from human written text. |
| Approach: | They propose a neural network detector that detects manipulated news articles by reasoning about the facts mentioned in the article. |
| Outcome: | The proposed detector outperforms the state-of-the-art detector in accuracy. |
Copied to clipboard
| Challenge: | AfroLID is a neural LID toolkit for 517 African languages and varieties. |
| Approach: | They propose to exploit a multi-domain web dataset manually curated from across 14 language families utilizing five orthographic systems to exploit AfroLID. |
| Outcome: | The proposed tool outperforms existing tools on the acutely under-served Twitter domain. |
Copied to clipboard
| Challenge: | Existing work on dialect prediction is limited to coarse-grained varieties . a new language model, MARBERT, can predict micro-dialects with 9.9% F1, 76 better than a majority class baseline. |
| Approach: | They propose a new task of Micro-Dialect Identification (MDI) that can predict a fine-grained variety given a single message. |
| Outcome: | The proposed model predicts micro-dialects with 9.9% F1, 76 better than a majority class baseline. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have diverse applications, encompassing both open-ended tasks (e.g., brainstorming and chat) and closed-ended ones (eg. question answering). |
| Approach: | They construct PP prompts for Large Language Models (LLMs) that estimate the performance of specific deep neural network architectures on downstream tasks. |
| Outcome: | The proposed model achieves a SoTA mean absolute error and a slight degradation in rank correlation coefficient compared to baseline predictors in machine translation tasks. |
Copied to clipboard
| Challenge: | Recent work on distilling Whisper’s knowledge into small models using pseudo-labels shows promising performance while reducing the size by up to 50%. |
| Approach: | They propose a framework that distills Whisper’s knowledge into small models using pseudo-labels and reduces the size by up to 50%. |
| Outcome: | The proposed model outperforms the teacher model by 5-7 WER points and is 25-50% more efficient when scaling the data. |
Copied to clipboard
| Challenge: | Current text generative models excel in producing text that matches the style of human language reasonably well. |
| Approach: | They conduct an in-depth error analysis of the state-of-the-art detector and discuss research directions to guide future work in this exciting area. |
| Outcome: | The proposed detectors can distinguish between human and text generated by the model and can be used to generate fake news and fake product reviews. |
Copied to clipboard
| Challenge: | Existing models produce outputs that are too advanced or vague for younger learners and there are no standardized benchmarks to evaluate their ability to adapt across cognitive and developmental stages. |
| Approach: | They propose to use a benchmark to assess LLMs' ability to adapt to different grade levels and to use it to evaluate their model's performance. |
| Outcome: | The proposed framework assesses the ability of large language models to adapt to grade levels across a range of subjects and grades. |
Copied to clipboard
| Challenge: | generative pretraining (GPT) scholarship remains acutely anglocentric, leaving serious gaps in our understanding of the whole class of autoregressive models. |
| Approach: | They propose to use Arabic autoregressive models to evaluate their performance . they use a benchmark to evaluate the models and code for experimenting with them . |
| Outcome: | JASMINE is a suite of powerful Arabic autoregressive models . it shows powerful performance intrinsically and in few-shot learning on a wide range of tasks. |