Papers by Hamdy Mubarak

22 papers
From RAG to Agentic RAG for Faithful Islamic Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used for Islamic question answering, where ungrounded responses may carry serious religious consequences.
Approach: They propose a bilingual, bilingual, Arabic/English benchmark with atomic single-gold answers that measures hallucination and abstention.
Outcome: The proposed model improves accuracy and robustness even with a small model.
Part-of-Speech Tagging for Arabic Gulf Dialect Using Bi-LSTM (L18-1)

Copied to clipboard

Challenge: Part-of-speech (POS) tagging is one of the most important building blocks in many natural language processing (NLP) applications.
Approach: They propose to use a POS tagger for Arabic Gulf dialect to improve POS tagging accuracy.
Outcome: The proposed POS tagger improves POS tagging accuracy for the Arabic Gulf dialect from 75% accuracy to 91% accuracy using a bi-LSTM labeler.
ASAD: Arabic Social media Analytics and unDerstanding (2021.eacl-demos)

Copied to clipboard

Challenge: Currently, there are no publicly available tools for analyzing Arabic social media, such as ADIDA and CAMeL, which are not trained with Twitter data.
Approach: They propose to use Arabic social media analysis and unDerstanding to analyze tweets using a web API and a user interface.
Outcome: The proposed system allows users to determine dialects, sentiment, news category, offensiveness, hate speech, adult content, and spam in Arabic tweets.
ArCovidVac: Analyzing Arabic Tweets About COVID-19 Vaccination (2022.lrec-1)

Copied to clipboard

Challenge: Social media are integrated with our daily life and are used to circulate information.
Approach: They develop and publicly release the first largest manually annotated Arabic tweet dataset for COVID-19 vaccination campaign.
Outcome: The proposed dataset is the largest manually annotated Arabic tweet dataset for COVID-19 vaccination campaign, covering many countries in the Arab region.
QASR: QCRI Aljazeera Speech Resource A Large Scale Annotated Arabic Speech Corpus (2021.acl-long)

Copied to clipboard

Challenge: QASR is the largest transcribed Arabic speech corpus in the broadcast domain.
Approach: They introduce the largest transcribed Arabic speech corpus, QASR, collected from the broadcast domain.
Outcome: The proposed dataset contains 2,000 hours of speech sampled at 16kHz crawled from Aljazeera news channel.
Fanar-Sadiq: A Multi-Agent Architecture for Grounded Islamic QA (2026.acl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) can answer religious knowledge queries fluently, but they often hallucinate and misattribute sources.
Approach: They propose a bilingual Arabic-English Islamic QA system that uses a multi-agent, tool-augmented architecture to route Islamic queries to specialized modules.
Outcome: The proposed system is based on a multi-agent, tool-augmented architecture and has received over 1.9M accesses in less than a year.
LAraBench: Benchmarking Arabic AI with Large Language Models (2024.eacl-long)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have significantly influenced the landscape of language and speech research.
Approach: They used GPT-3.5-turbo, GPT-4, BLOOMZ, Jais-13b-chat, Whisper, and USM to tackle 33 distinct tasks across 61 datasets.
Outcome: The proposed model outperforms SOTA models in zero-shot learning, with a few exceptions.
A Multi-Platform Arabic News Comment Dataset for Offensive Language Detection (2020.lrec-1)

Copied to clipboard

Challenge: Social media platforms allow users to engage in conversation with limited accountability, causing hate crimes and mental harm to targeted individuals.
Approach: They propose to make public a new dialectal Arabic news comment dataset . they analyze distinctive lexical content along with the use of emojis in offensive comments .
Outcome: The proposed dataset analyzes offensive language and distinctive lexical content along with the use of emojis on Twitter, Facebook, and YouTube.
Arabic Curriculum Analysis (2020.coling-demos)

Copied to clipboard

Challenge: Effective language curricula are critical to teaching communication skills . a platform that analyzes curricular content can help identify shortcomings .
Approach: They propose a platform that analyzes Arabic curricula and provides insights into their content.
Outcome: The proposed system analyzes Arabic curricula and provides insights into their content . it provides statistics about word usage and morphological forms in different grades .
Multi-Dialect Arabic POS Tagging: A CRF Approach (L18-1)

Copied to clipboard

Challenge: Existing work on dialectal POS tagging is rather scant with POS tags for most dialects being nonexistent or of limited availability.
Approach: They propose a dataset of POS-tagged Arabic tweets in four major dialects and a tagging guideline for each dialect.
Outcome: The proposed model can tag four different dialects with an average accuracy of 89.3%.
A System for Diacritizing Four Varieties of Arabic (D19-3)

Copied to clipboard

Challenge: Short vowels, aka diacritics, are omitted when writing different varieties of Arabic . diacritization is essential for language learning and text-to-speech applications .
Approach: They propose a system for recovering diacritics in Arabic without short vowels . they use a character-based sequence-to-sequence deep learning model .
Outcome: The proposed system beats all previous SOTA systems for Arabic varieties . it uses a character-based sequence-to-sequence deep learning model .
AraSafe: Benchmarking Safety in Arabic LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: AraSafe is the first large-scale native Arabic safety benchmark for large language models (LLMs) it addresses the pressing need for culturally and linguistically representative evaluation resources.
Approach: They propose to use Arabic prompts to annotate harmful and non-harmful prompts into nine fine-grained safety categories to support classifiers for harmful content.
Outcome: The proposed benchmarks address the need for culturally and linguistically representative evaluation resources.
Advancing Arabic Diacritization: Improved Datasets, Benchmarking, and State-of-the-Art Models (2025.emnlp-main)

Copied to clipboard

Challenge: Arabic diacritics are typically omitted in written Arabic, leading to ambiguity . authors propose a methodology to analyze and refine a large diacritized corpus .
Approach: They propose a methodology to analyze and refine a large diacritized corpus to improve training quality.
Outcome: The proposed model achieves state-of-the-art results with 3.12% and 2.70% WER on WikiNews-2014 and Wikinews-2024.
Build Fast and Accurate Lemmatization for Arabic (L18-1)

Copied to clipboard

Challenge: Lemmatization is the process of finding the base form (lemma) of a word by considering its inflected forms.
Approach: They propose a lemmatizer for Arabic with a dataset that can be used to test lemma accuracy.
Outcome: The proposed algorithm outperforms state-of-the-art Arabic lemmatization in accuracy and speed.
LLMeBench: A Flexible Framework for Accelerating LLMs Benchmarking (2024.eacl-demo)

Copied to clipboard

Challenge: Recent development and success of Large Language Models necessitate evaluation of their performance across diverse NLP tasks in different languages.
Approach: They propose a framework that can be customized to evaluate LLMs for any NLP task, regardless of language.
Outcome: The LLMeBench framework can be customized to evaluate LLMs for any NLP task, regardless of language.
Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification (2024.findings-acl)

Copied to clipboard

Challenge: Large language models are notorious for producing erroneous claims in their output.
Approach: They propose a fact-checking and hallucination detection pipeline based on token-level uncertainty quantification that removes the impact of uncertainty about what claim to generate on the current step and what surface form to use.
Outcome: The proposed method can fact-check the atomic claims in the output of large language models.
Beyond Orthography: Automatic Recovery of Short Vowels and Dialectal Sounds in Arabic (2024.acl-long)

Copied to clipboard

Challenge: Existing algorithms for recognizing borrowed and dialectal sounds are limited to Arabic, a dialect-rich language containing more than 22 major dialects.
Approach: They propose a framework to recognize borrowed and dialectal sounds within phonologically diverse and dialect-rich languages that extends beyond its standard orthographic sound sets.
Outcome: The proposed framework improves character error rate by 7% with only one and half hours of training data compared to the baseline.
So Hateful! Building a Multi-Label Hate Speech Annotated Arabic Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Social media enables widespread propagation of hate speech targeting groups based on ethnicity, religion, or other characteristics.
Approach: They analyze 70,000 Arabic tweets to identify hate speech patterns and train models . 15% of tweets contain offensive language while 6% have hate speech . authors hope to prevent spread of hateful content on social media platforms .
Outcome: The analysis of 70,000 Arabic tweets shows that 15% of tweets contain offensive language while 6% have hate speech . 10% of tweet provide verifiable factual claims, and 7% are deemed important .
Nahw: A Comprehensive Benchmark of Arabic Grammar Understanding, Error Detection, Correction, and Explanation (2026.eacl-long)

Copied to clipboard

Challenge: Existing corpora address individual linguistic aspects like spelling or diacritization, but rarely provide explanations of grammatical errors. Existing datasets and benchmarks that capture Arabic's grammatological complexity are scarce.
Approach: They propose a benchmark for Arabic grammar that covers error detection, correction, and explanation.
Outcome: The proposed model performs better on GPT-4o than on the best performing model (ALLaM-7B) despite fine tuning with synthetic data, the model perform better on Arabic grammar tasks.
Halwasa: Quantify and Analyze Hallucinations in Large Language Models: Arabic as a Case Study (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) generate text that is factually incorrect, nonsensical, or misleading.
Approach: They create a large Arabic dataset that contains 10K of LLM generated sentences and annotate it for factuality and correctness.
Outcome: The proposed dataset analyzes 10K of generated sentences and finds 25% of them are factually incorrect.
Highly Effective Arabic Diacritization using Sequence to Sequence Modeling (N19-1)

Copied to clipboard

Challenge: Arabic text is written without short vowels (or diacritics) their presence is essential for properly verbalizing Arabic .
Approach: They propose a character-level sequence-to-sequence deep learning model that recovers both types of diacritics without the use of explicit feature engineering.
Outcome: The proposed model outperforms all previous state-of-the-art models on overlapping windows of words . it achieves a word error rate (WER) of 4.49% compared to the state- of-the art systems .
Fighting the COVID-19 Infodemic: Modeling the Perspective of Journalists, Fact-Checkers, Social Media Platforms, Policy Makers, and the Society (2021.findings-emnlp)

Copied to clipboard

Challenge: a dataset of 16K manually annotated tweets is used to analyze disinformation . the democratic nature of social media has raised questions about the quality and the factuality of the information that is shared on these platforms.
Approach: They use a dataset of manually annotated tweets to analyze COVID-19 disinformation . they show that tweets contain fake cures, rumors, conspiracy theories and xenophobia .
Outcome: The proposed dataset shows that it is useful in monolingual vs. multilingual settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations