Papers by Mustafa Jarrar

12 papers
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset (2025.findings-emnlp)

Copied to clipboard

Challenge: Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets.
Approach: They propose to construct a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding.
Outcome: The proposed dataset covers ten culturally significant domains covering all Arab countries and includes two evaluation benchmarks (PEARL and PEARL-LITE) and a specialized subset (PearL-X).
From RAG to Agentic RAG for Faithful Islamic Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used for Islamic question answering, where ungrounded responses may carry serious religious consequences.
Approach: They propose a bilingual, bilingual, Arabic/English benchmark with atomic single-gold answers that measures hallucination and abstention.
Outcome: The proposed model improves accuracy and robustness even with a small model.
Active Learning for Multidialectal Arabic POS Tagging (2025.findings-emnlp)

Copied to clipboard

Challenge: Multidialectal Arabic POS tagging is challenging due to the morphological richness and high variability among dialects.
Approach: They propose an active learning approach for multidialectal Arabic POS tagging . they annotate approximately 15,000 tokens, reducing the annotation requirement by about 2,000 tokens .
Outcome: The proposed approach achieves 97.6% accuracy on the Emirati corpus.
Qabas: An Open-Source Arabic Lexicographic Database (2024.lrec-main)

Copied to clipboard

Challenge: Qabas is an open-source Arabic lexicon designed for NLP applications.
Approach: They propose to link lemmas from 110 lexicons into a morphologically annotated Arabic lexicoma.
Outcome: Qabas lexical entries (lemmas) are assembled by linking lemmas from 110 lexicons.
Konooz: Multi-domain Multi-dialect Corpus for Named Entity Recognition (2025.findings-acl)

Copied to clipboard

Challenge: Using the Wojood framework, we compare existing Arabic Named Entity Recognition models with domain and dialect divergence and resource scarcity.
Approach: They propose a multi-dimensional Arabic named entity corpus covering 16 dialects across 10 domains and an annotation scheme using the Wojood guidelines.
Outcome: The proposed model performs better on 16 dialects across 10 domains and 16 domains, while other models struggle with different dialects and domains.
Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs (2026.acl-long)

Copied to clipboard

Challenge: Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than modern standard Arabic (MSA).
Approach: They propose a large-scale, community-driven, human-translated dataset to bridge this gap . Alexandria covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata .
Outcome: The Alexandria dataset covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . Alexandria is a training resource and a rigorous benchmark for evaluating MT and LLMs based on the Alexandria dataset .
WojoodRelations: Arabic Relation Extraction Corpus and Modeling (2025.emnlp-main)

Copied to clipboard

Challenge: Existing work on Arabic RE remains limited due to the language’s rich morphology and syntactic complexity, and the lack of large, high-quality datasets.
Approach: They propose to use WojoodRelations to extract relation relationships from Arabic textual data using relation-aware templates and GPT-Joint to perform relation-based retrieval.
Outcome: The proposed method achieves a Cohen’s of 0.92, indicating high reliability, and supervised models achieve 92.89% F1 for RE, while LLMs obtain 72.73% F1 .
Wojood: Nested Arabic Named Entity Corpus and Recognition using BERT (2022.lrec-1)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is integral to many NLP applications such as chatbots and question answering.
Approach: They propose to annotate Arabic nested entities instead of flat annotations by manually annotating 550K tokens with 21 entity types including person, organization, location, event and date.
Outcome: The proposed model achieved an overall micro F1-score of 0.884 and the annotation guidelines and source code are publicly available.
AdabNER: Arabic Digital Archive Books with Nested Entity Recognition (2026.acl-long)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a subtask of information extraction that classifies entities into predefined categories like person names.
Approach: They propose a large-scale nested Arabic Named Entity Recognition dataset . they fine-tuned five pre-trained Arabic BERT encoders in two settings .
Outcome: The first large-scale nested NER dataset for Arabic literary texts is published online . the dataset yields 78,530 entity mentions, 18.96% of which are nestated .
Curras + Baladi: Towards a Levantine Corpus (2022.lrec-1)

Copied to clipboard

Challenge: The processing of the Arabic language is a complex field of research due to the complex and rich morphology of Arabic, its high degree of ambiguity, and the presence of several regional varieties that need to be processed while taking into account their unique characteristics.
Approach: They propose to revise the Palestinian morphologically annotated corpus and a Lebanese corpus to bridge nuanced linguistic gaps between the two highly mutually intelligible dialects.
Outcome: The revised corpus can be used as a more general Levantine corpus.
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition (2024.emnlp-main)

Copied to clipboard

Challenge: despite recent advances in speech processing, the majority of world languages and dialects remain uncovered.
Approach: They propose to collect and transcribe a new Arabic dataset for eight dialects . they also develop strong baselines exploiting the new dataset .
Outcome: The proposed dataset covers eight Arabic dialects, including Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations