Papers by Hamdy Mubarak
From RAG to Agentic RAG for Faithful Islamic Question Answering (2026.findings-acl)
Copied to clipboard
Gagan Bhatia, Hamdy Mubarak, Mustafa Jarrar, George Mikros, Fadi Zaraket, Mahmoud Alhirthani, Mutaz al-Khatib, Logan Cochrane, Kareem Mohamed Darwish, Rashid Yahiaoui, Firoj Alam
| Challenge: | Large Language Models (LLMs) are increasingly used for Islamic question answering, where ungrounded responses may carry serious religious consequences. |
| Approach: | They propose a bilingual, bilingual, Arabic/English benchmark with atomic single-gold answers that measures hallucination and abstention. |
| Outcome: | The proposed model improves accuracy and robustness even with a small model. |
Part-of-Speech Tagging for Arabic Gulf Dialect Using Bi-LSTM (L18-1)
Copied to clipboard
| Challenge: | Part-of-speech (POS) tagging is one of the most important building blocks in many natural language processing (NLP) applications. |
| Approach: | They propose to use a POS tagger for Arabic Gulf dialect to improve POS tagging accuracy. |
| Outcome: | The proposed POS tagger improves POS tagging accuracy for the Arabic Gulf dialect from 75% accuracy to 91% accuracy using a bi-LSTM labeler. |
ASAD: Arabic Social media Analytics and unDerstanding (2021.eacl-demos)
Copied to clipboard
| Challenge: | Currently, there are no publicly available tools for analyzing Arabic social media, such as ADIDA and CAMeL, which are not trained with Twitter data. |
| Approach: | They propose to use Arabic social media analysis and unDerstanding to analyze tweets using a web API and a user interface. |
| Outcome: | The proposed system allows users to determine dialects, sentiment, news category, offensiveness, hate speech, adult content, and spam in Arabic tweets. |
ArCovidVac: Analyzing Arabic Tweets About COVID-19 Vaccination (2022.lrec-1)
Copied to clipboard
| Challenge: | Social media are integrated with our daily life and are used to circulate information. |
| Approach: | They develop and publicly release the first largest manually annotated Arabic tweet dataset for COVID-19 vaccination campaign. |
| Outcome: | The proposed dataset is the largest manually annotated Arabic tweet dataset for COVID-19 vaccination campaign, covering many countries in the Arab region. |
QASR: QCRI Aljazeera Speech Resource A Large Scale Annotated Arabic Speech Corpus (2021.acl-long)
Copied to clipboard
| Challenge: | QASR is the largest transcribed Arabic speech corpus in the broadcast domain. |
| Approach: | They introduce the largest transcribed Arabic speech corpus, QASR, collected from the broadcast domain. |
| Outcome: | The proposed dataset contains 2,000 hours of speech sampled at 16kHz crawled from Aljazeera news channel. |
Fanar-Sadiq: A Multi-Agent Architecture for Grounded Islamic QA (2026.acl-industry)
Copied to clipboard
Ummar Abbas, Mourad Ouzzani, Mohamed Y. Eltabakh, Omar Sinan, Gagan Bhatia, Hamdy Mubarak, Majd Hawasly, Mohammed Qusay Hashim, Kareem Mohamed Darwish, Firoj Alam
| Challenge: | Large language models (LLMs) can answer religious knowledge queries fluently, but they often hallucinate and misattribute sources. |
| Approach: | They propose a bilingual Arabic-English Islamic QA system that uses a multi-agent, tool-augmented architecture to route Islamic queries to specialized modules. |
| Outcome: | The proposed system is based on a multi-agent, tool-augmented architecture and has received over 1.9M accesses in less than a year. |
LAraBench: Benchmarking Arabic AI with Large Language Models (2024.eacl-long)
Copied to clipboard
Ahmed Abdelali, Hamdy Mubarak, Shammur Chowdhury, Maram Hasanain, Basel Mousi, Sabri Boughorbel, Samir Abdaljalil, Yassine El Kheir, Daniel Izham, Fahim Dalvi, Majd Hawasly, Nizi Nazar, Youssef Elshahawy, Ahmed Ali, Nadir Durrani, Natasa Milic-Frayling, Firoj Alam
| Challenge: | Recent advances in Large Language Models (LLMs) have significantly influenced the landscape of language and speech research. |
| Approach: | They used GPT-3.5-turbo, GPT-4, BLOOMZ, Jais-13b-chat, Whisper, and USM to tackle 33 distinct tasks across 61 datasets. |
| Outcome: | The proposed model outperforms SOTA models in zero-shot learning, with a few exceptions. |
A Multi-Platform Arabic News Comment Dataset for Offensive Language Detection (2020.lrec-1)
Copied to clipboard
Shammur Absar Chowdhury, Hamdy Mubarak, Ahmed Abdelali, Soon-gyo Jung, Bernard J. Jansen, Joni Salminen
| Challenge: | Social media platforms allow users to engage in conversation with limited accountability, causing hate crimes and mental harm to targeted individuals. |
| Approach: | They propose to make public a new dialectal Arabic news comment dataset . they analyze distinctive lexical content along with the use of emojis in offensive comments . |
| Outcome: | The proposed dataset analyzes offensive language and distinctive lexical content along with the use of emojis on Twitter, Facebook, and YouTube. |
Arabic Curriculum Analysis (2020.coling-demos)
Copied to clipboard
| Challenge: | Effective language curricula are critical to teaching communication skills . a platform that analyzes curricular content can help identify shortcomings . |
| Approach: | They propose a platform that analyzes Arabic curricula and provides insights into their content. |
| Outcome: | The proposed system analyzes Arabic curricula and provides insights into their content . it provides statistics about word usage and morphological forms in different grades . |
Multi-Dialect Arabic POS Tagging: A CRF Approach (L18-1)
Copied to clipboard
Kareem Darwish, Hamdy Mubarak, Ahmed Abdelali, Mohamed Eldesouki, Younes Samih, Randah Alharbi, Mohammed Attia, Walid Magdy, Laura Kallmeyer
| Challenge: | Existing work on dialectal POS tagging is rather scant with POS tags for most dialects being nonexistent or of limited availability. |
| Approach: | They propose a dataset of POS-tagged Arabic tweets in four major dialects and a tagging guideline for each dialect. |
| Outcome: | The proposed model can tag four different dialects with an average accuracy of 89.3%. |
A System for Diacritizing Four Varieties of Arabic (D19-3)
Copied to clipboard
| Challenge: | Short vowels, aka diacritics, are omitted when writing different varieties of Arabic . diacritization is essential for language learning and text-to-speech applications . |
| Approach: | They propose a system for recovering diacritics in Arabic without short vowels . they use a character-based sequence-to-sequence deep learning model . |
| Outcome: | The proposed system beats all previous SOTA systems for Arabic varieties . it uses a character-based sequence-to-sequence deep learning model . |
AraSafe: Benchmarking Safety in Arabic LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | AraSafe is the first large-scale native Arabic safety benchmark for large language models (LLMs) it addresses the pressing need for culturally and linguistically representative evaluation resources. |
| Approach: | They propose to use Arabic prompts to annotate harmful and non-harmful prompts into nine fine-grained safety categories to support classifiers for harmful content. |
| Outcome: | The proposed benchmarks address the need for culturally and linguistically representative evaluation resources. |
Advancing Arabic Diacritization: Improved Datasets, Benchmarking, and State-of-the-Art Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Arabic diacritics are typically omitted in written Arabic, leading to ambiguity . authors propose a methodology to analyze and refine a large diacritized corpus . |
| Approach: | They propose a methodology to analyze and refine a large diacritized corpus to improve training quality. |
| Outcome: | The proposed model achieves state-of-the-art results with 3.12% and 2.70% WER on WikiNews-2014 and Wikinews-2024. |
Build Fast and Accurate Lemmatization for Arabic (L18-1)
Copied to clipboard
| Challenge: | Lemmatization is the process of finding the base form (lemma) of a word by considering its inflected forms. |
| Approach: | They propose a lemmatizer for Arabic with a dataset that can be used to test lemma accuracy. |
| Outcome: | The proposed algorithm outperforms state-of-the-art Arabic lemmatization in accuracy and speed. |
LLMeBench: A Flexible Framework for Accelerating LLMs Benchmarking (2024.eacl-demo)
Copied to clipboard
Fahim Dalvi, Maram Hasanain, Sabri Boughorbel, Basel Mousi, Samir Abdaljalil, Nizi Nazar, Ahmed Abdelali, Shammur Absar Chowdhury, Hamdy Mubarak, Ahmed Ali
| Challenge: | Recent development and success of Large Language Models necessitate evaluation of their performance across diverse NLP tasks in different languages. |
| Approach: | They propose a framework that can be customized to evaluate LLMs for any NLP task, regardless of language. |
| Outcome: | The LLMeBench framework can be customized to evaluate LLMs for any NLP task, regardless of language. |
Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification (2024.findings-acl)
Copied to clipboard
Ekaterina Fadeeva, Aleksandr Rubashevskii, Artem Shelmanov, Sergey Petrakov, Haonan Li, Hamdy Mubarak, Evgenii Tsymbalov, Gleb Kuzmin, Alexander Panchenko, Timothy Baldwin, Preslav Nakov, Maxim Panov
| Challenge: | Large language models are notorious for producing erroneous claims in their output. |
| Approach: | They propose a fact-checking and hallucination detection pipeline based on token-level uncertainty quantification that removes the impact of uncertainty about what claim to generate on the current step and what surface form to use. |
| Outcome: | The proposed method can fact-check the atomic claims in the output of large language models. |
Beyond Orthography: Automatic Recovery of Short Vowels and Dialectal Sounds in Arabic (2024.acl-long)
Copied to clipboard
| Challenge: | Existing algorithms for recognizing borrowed and dialectal sounds are limited to Arabic, a dialect-rich language containing more than 22 major dialects. |
| Approach: | They propose a framework to recognize borrowed and dialectal sounds within phonologically diverse and dialect-rich languages that extends beyond its standard orthographic sound sets. |
| Outcome: | The proposed framework improves character error rate by 7% with only one and half hours of training data compared to the baseline. |
So Hateful! Building a Multi-Label Hate Speech Annotated Arabic Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | Social media enables widespread propagation of hate speech targeting groups based on ethnicity, religion, or other characteristics. |
| Approach: | They analyze 70,000 Arabic tweets to identify hate speech patterns and train models . 15% of tweets contain offensive language while 6% have hate speech . authors hope to prevent spread of hateful content on social media platforms . |
| Outcome: | The analysis of 70,000 Arabic tweets shows that 15% of tweets contain offensive language while 6% have hate speech . 10% of tweet provide verifiable factual claims, and 7% are deemed important . |
Nahw: A Comprehensive Benchmark of Arabic Grammar Understanding, Error Detection, Correction, and Explanation (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing corpora address individual linguistic aspects like spelling or diacritization, but rarely provide explanations of grammatical errors. Existing datasets and benchmarks that capture Arabic's grammatological complexity are scarce. |
| Approach: | They propose a benchmark for Arabic grammar that covers error detection, correction, and explanation. |
| Outcome: | The proposed model performs better on GPT-4o than on the best performing model (ALLaM-7B) despite fine tuning with synthetic data, the model perform better on Arabic grammar tasks. |
Halwasa: Quantify and Analyze Hallucinations in Large Language Models: Arabic as a Case Study (2024.lrec-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) generate text that is factually incorrect, nonsensical, or misleading. |
| Approach: | They create a large Arabic dataset that contains 10K of LLM generated sentences and annotate it for factuality and correctness. |
| Outcome: | The proposed dataset analyzes 10K of generated sentences and finds 25% of them are factually incorrect. |
Highly Effective Arabic Diacritization using Sequence to Sequence Modeling (N19-1)
Copied to clipboard
| Challenge: | Arabic text is written without short vowels (or diacritics) their presence is essential for properly verbalizing Arabic . |
| Approach: | They propose a character-level sequence-to-sequence deep learning model that recovers both types of diacritics without the use of explicit feature engineering. |
| Outcome: | The proposed model outperforms all previous state-of-the-art models on overlapping windows of words . it achieves a word error rate (WER) of 4.49% compared to the state- of-the art systems . |
Fighting the COVID-19 Infodemic: Modeling the Perspective of Journalists, Fact-Checkers, Social Media Platforms, Policy Makers, and the Society (2021.findings-emnlp)
Copied to clipboard
Firoj Alam, Shaden Shaar, Fahim Dalvi, Hassan Sajjad, Alex Nikolov, Hamdy Mubarak, Giovanni Da San Martino, Ahmed Abdelali, Nadir Durrani, Kareem Darwish, Abdulaziz Al-Homaid, Wajdi Zaghouani, Tommaso Caselli, Gijs Danoe, Friso Stolk, Britt Bruntink, Preslav Nakov
| Challenge: | a dataset of 16K manually annotated tweets is used to analyze disinformation . the democratic nature of social media has raised questions about the quality and the factuality of the information that is shared on these platforms. |
| Approach: | They use a dataset of manually annotated tweets to analyze COVID-19 disinformation . they show that tweets contain fake cures, rumors, conspiracy theories and xenophobia . |
| Outcome: | The proposed dataset shows that it is useful in monolingual vs. multilingual settings. |