Papers by Walid Magdy
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset (2025.findings-emnlp)
Copied to clipboard
Fakhraddin Alwajih, Samar M. Magdy, Abdellah El Mekki, Omer Nacar, Youssef Nafea, Safaa Taher Abdelfadil, Abdulfattah Mohammed Yahya, Hamzah Luqman, Nada Almarwani, Samah Aloufi, Baraah Qawasmeh, Houdaifa Atou, Serry Sibaee, Hamzah A. Alsayadi, Walid Al-Dhabyani, Maged S. Al-shaibani, Aya El aatar, Nour Qandos, Rahaf Alhamouri, Samar Ahmad, Mohammed Anwar AL-Ghrawi, Aminetou Yacoub, Ruwa AbuHweidi, Vatimetou Mohamed Lemin, Reem Abdel-Salam, Ahlam Bashiti, Adel Ammar, Aisha Alansari, Ahmed Ashraf, Nora Alturayeif, Alcides Alcoba Inciarte, AbdelRahim A. Elmadany, Mohamedou Cheikh Tourad, Ismail Berrada, Mustafa Jarrar, Shady Shehata, Muhammad Abdul-Mageed
| Challenge: | Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. |
| Approach: | They propose to construct a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding. |
| Outcome: | The proposed dataset covers ten culturally significant domains covering all Arab countries and includes two evaluation benchmarks (PEARL and PEARL-LITE) and a specialized subset (PearL-X). |
DLAMA: A Framework for Curating Culturally Diverse Facts for Probing the Knowledge of Pretrained Language Models (2023.findings-acl)
Copied to clipboard
| Challenge: | a few benchmarking datasets have been released to evaluate the factual knowledge of pretrained language models. |
| Approach: | They propose a framework for curating factual triples from Wikidata that are culturally diverse. |
| Outcome: | The proposed framework is built of factual triples from three pairs of contrasting cultures with 78,259 triples. |
Culture Matters in Toxic Language Detection in Persian (2025.acl-long)
Copied to clipboard
| Challenge: | Toxic language detection is crucial for creating safer online environments and limiting the spread of harmful content. |
| Approach: | They compare different methods for toxic language detection in Persian to fine-tune, enrich data, and cross-lingual transfer learning. |
| Outcome: | The language of a country with cultural similarities to Persian yields better results in transfer learning. |
Urban Dictionary Embeddings for Slang NLP Applications (2020.lrec-1)
Copied to clipboard
| Challenge: | a new set of word embeddings is released to improve word embedment performance . word embeds provide useful representations of meanings of words in vectors . |
| Approach: | They present a set of word embeddings trained on Urban Dictionary . they show they have high performance across a range of common word embeding evaluations . |
| Outcome: | The first set of word embeddings trained on Urban Dictionary has high performance . the embeddables perform better on a range of common word evaluation tasks . |
Part-of-Speech Tagging for Arabic Gulf Dialect Using Bi-LSTM (L18-1)
Copied to clipboard
| Challenge: | Part-of-speech (POS) tagging is one of the most important building blocks in many natural language processing (NLP) applications. |
| Approach: | They propose to use a POS tagger for Arabic Gulf dialect to improve POS tagging accuracy. |
| Outcome: | The proposed POS tagger improves POS tagging accuracy for the Arabic Gulf dialect from 75% accuracy to 91% accuracy using a bi-LSTM labeler. |
Exploring Author Context for Detecting Intended vs Perceived Sarcasm (P19-1)
Copied to clipboard
| Challenge: | Existing studies on textual sarcasm detection use manual labelling and tag-based distant supervision to detect sarcasm. |
| Approach: | They define author context as the embedded representation of their historical tweets and suggest neural models that extract these representations. |
| Outcome: | The proposed models achieve state-of-the-art on two datasets labelled manually and via tag-based distant supervision indicating a difference between intended and perceived sarcasm . |
AX-MABSA: A Framework for Extremely Weakly Supervised Multi-label Aspect Based Sentiment Analysis (2022.emnlp-main)
Copied to clipboard
| Challenge: | Aspect Based Sentiment Analysis is a dominant research area with potential applications in social media analytics, business, finance, and health. |
| Approach: | They propose a weakly supervised multi-label Aspect Category Sentiment Analysis framework which does not use any labelled data. |
| Outcome: | The proposed framework outperforms weakly supervised baselines on four benchmark datasets and is able to generate multiple aspect category-sentiment pairs per review sentence. |
Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs (2026.acl-long)
Copied to clipboard
Abdellah EL Mekki, Samar M. Magdy, Houdaifa Atou, Ruwa AbuHweidi, Baraah Qawasmeh, Omer Nacar, Thikra Al-hibiri, Razan Saadie, Hamzah A. Alsayadi, Nadia Ghezaiel Hammouda, Alshima Mohammed Alkhazimi, Aya Hamod, Al-Yas Yaqoob Al-Ghafri, Wesam El-Sayed, Asila Ismail al Sharji, Mohamad Ballout, Anas Belfathi, Karim Ghaddar, Serry Sibaee, Alaa Aoun, Aeej Mohammed Aseri, Lina Abureesh, Ahlam Bashiti, Majdal Yousef, Abdulaziz Hafiz, Yehdih Mohamed, Emira Hamedtou, Brakehe Emehah, Rahaf Alhamouri, Youssef Nafea, Aya El Aatar, Walid Al-Dhabyani, Emhemed S. Hamed, Sara Shatnawi, Fakhraddin Alwajih, Khalid Elkhidir, Ashwag Alasmari, Abdurrahman Gerrio, Omar Said Alshahri, AbdelRahim A. Elmadany, Ismail Berrada, Amir Azad Adli Al-kathiri, Fadi Zaraket, Mustafa Jarrar, Yahya Mohamed EL Hadj, Hassan Alhuzali, Muhammad Abdul-Mageed
| Challenge: | Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than modern standard Arabic (MSA). |
| Approach: | They propose a large-scale, community-driven, human-translated dataset to bridge this gap . Alexandria covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . |
| Outcome: | The Alexandria dataset covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . Alexandria is a training resource and a rigorous benchmark for evaluating MT and LLMs based on the Alexandria dataset . |
iSarcasm: A Dataset of Intended Sarcasm (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods for detecting intended sarcasm have shown low performance compared to previous studies. |
| Approach: | They propose a dataset of tweets labeled for intended sarcasm by their authors . they aim to encourage future NLP research to develop methods for detecting sarkasmus in text as intended by the authors of the text . |
| Outcome: | The proposed model shows that existing methods are biased or obvious and sarcasm could be understudied. |
Validating Automatic Evaluation of Controllable Counterspeech Generation: Rankings Matter More Than Scores (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing methods for evaluating attributes of counterspeech are limited and the validity of such evaluations is questionable when the classifiers themselves have only modest performance. |
| Approach: | They examine the automatic evaluation of counterspeech attributes using a multi-attribute counterseech dataset containing 2,728 samples. |
| Outcome: | The proposed model can be trusted by classifier validation, and it can rank models with confidence. |
Revisiting Common Assumptions about Arabic Dialects in NLP (2025.acl-long)
Copied to clipboard
| Challenge: | Existing assumptions about Arabic dialect variation are not quantitatively verified. |
| Approach: | They extend and analyze Arabic dialects to assess their validity using a multi-label dataset . they find that the assumptions oversimplify reality and are not always accurate . |
| Outcome: | The proposed methods oversimplify reality and are not always accurate, the authors argue . they show that the proposed assumptions oversimply represent reality and may hinder future work . |
Multi-Dialect Arabic POS Tagging: A CRF Approach (L18-1)
Copied to clipboard
Kareem Darwish, Hamdy Mubarak, Ahmed Abdelali, Mohamed Eldesouki, Younes Samih, Randah Alharbi, Mohammed Attia, Walid Magdy, Laura Kallmeyer
| Challenge: | Existing work on dialectal POS tagging is rather scant with POS tags for most dialects being nonexistent or of limited availability. |
| Approach: | They propose a dataset of POS-tagged Arabic tweets in four major dialects and a tagging guideline for each dialect. |
| Outcome: | The proposed model can tag four different dialects with an average accuracy of 89.3%. |
Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs (2025.acl-long)
Copied to clipboard
Fakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy, AbdelRahim A. Elmadany, Omer Nacar, El Moatez Billah Nagoudi, Reem Abdel-Salam, Hanin Atwany, Youssef Nafea, Abdulfattah Mohammed Yahya, Rahaf Alhamouri, Hamzah A. Alsayadi, Hiba Zayed, Sara Shatnawi, Serry Sibaee, Yasir Ech-chammakhy, Walid Al-Dhabyani, Marwa Mohamed Ali, Imen Jarraya, Ahmed Oumar El-Shangiti, Aisha Alraeesi, Mohammed Anwar AL-Ghrawi, Abdulrahman S. Al-Batati, Elgizouli Mohamed, Noha Taha Elgindi, Muhammed Saeed, Houdaifa Atou, Issam Ait Yahia, Abdelhak Bouayad, Mohammed Machrouh, Amal Makouar, Dania Alkawi, Mukhtar Mohamed, Safaa Taher Abdelfadil, Amine Ziad Ounnoughene, Anfel Rouabhia, Rwaa Assi, Ahmed Sorkatti, Mohamedou Cheikh Tourad, Anis Koubaa, Ismail Berrada, Mustafa Jarrar, Shady Shehata, Muhammad Abdul-Mageed
| Challenge: | a year-long community-driven project covering all 22 Arab countries evaluates the cultural and dialectal capabilities of large language models. |
| Approach: | They propose a project to evaluate the cultural and dialectal capabilities of large language models. |
| Outcome: | The project evaluates the cultural and dialectal capabilities of several frontier LLMs. |
Sarcasm Detection is Way Too Easy! An Empirical Comparison of Human and Machine Sarcasm Detection (2022.findings-emnlp)
Copied to clipboard
| Challenge: | sarcasm detection datasets focus on intended, rather than perceived sarcasm, but there is no comparison between human and machine performance. |
| Approach: | They collect author-annotated sarcasm datasets that focus on intended, rather than perceived sarcasticism . they compare human-level benchmarks to that of state-of-the-art sarkasmatic detection systems . |
| Outcome: | The proposed datasets compare human and machine performance on sarcastic tasks in English and Arabic. |
ALDi: Quantifying the Arabic Level of Dialectness of Text (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on Dialect Identification (DI) on the sentence level has focused on binary tasks, whereas ALDi treats the task as binary. |
| Approach: | They propose a dataset which contains 127,835 sentences manually labeled with their level of dialectness. |
| Outcome: | The proposed model can identify dialectness on a range of other corpora, providing a more nuanced picture than traditional DI systems. |
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition (2024.emnlp-main)
Copied to clipboard
Bashar Talafha, Karima Kadaoui, Samar Magdy, Mariem Habiboullah, Chafei Chafei, Ahmed El-Shangiti, Hiba Zayed, Mohamedou Tourad, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, Hour Mohamed, Fakhraddin Alwajih, Abdelrahman Mohamed, Abdellah El Mekki, El Moatez Billah Nagoudi, Benelhadj Saadia, Hamzah Alsayadi, Walid Al-Dhabyani, Sara Shatnawi, Yasir Ech-chammakhy, Amal Makouar, Yousra Berrachedi, Mustafa Jarrar, Shady Shehata, Ismail Berrada, Muhammad Abdul-Mageed
| Challenge: | despite recent advances in speech processing, the majority of world languages and dialects remain uncovered. |
| Approach: | They propose to collect and transcribe a new Arabic dataset for eight dialects . they also develop strong baselines exploiting the new dataset . |
| Outcome: | The proposed dataset covers eight Arabic dialects, including Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni. |
Chandler: An Explainable Sarcastic Response Generator (2021.emnlp-demo)
Copied to clipboard
| Challenge: | sarcasm generators assume intended meaning is opposite of literal meaning . sarcastically generated responses are more specific and coherent to input . |
| Approach: | They propose a system that generates sarcastic responses to a given utterance . they ground their generation process on a formal theory that unambiguously differentiates . |
| Outcome: | The proposed system generates sarcastic responses to a given utterance. |
Should a Chatbot be Sarcastic? Understanding User Preferences Towards Sarcasm Generation (2022.acl-long)
Copied to clipboard
| Challenge: | sarcasm generation research focused on creating more human-like interactions . previous research focused only on how to generate text that people perceive as sarkastic . |
| Approach: | They propose a theory-driven framework for generating sarcastic responses that allows us to control linguistic devices included during generation. |
| Outcome: | The proposed framework allows us to control the linguistic devices included during generation. |