Papers by Fadi Zaraket
From RAG to Agentic RAG for Faithful Islamic Question Answering (2026.findings-acl)
Copied to clipboard
Gagan Bhatia, Hamdy Mubarak, Mustafa Jarrar, George Mikros, Fadi Zaraket, Mahmoud Alhirthani, Mutaz al-Khatib, Logan Cochrane, Kareem Mohamed Darwish, Rashid Yahiaoui, Firoj Alam
| Challenge: | Large Language Models (LLMs) are increasingly used for Islamic question answering, where ungrounded responses may carry serious religious consequences. |
| Approach: | They propose a bilingual, bilingual, Arabic/English benchmark with atomic single-gold answers that measures hallucination and abstention. |
| Outcome: | The proposed model improves accuracy and robustness even with a small model. |
Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs (2026.acl-long)
Copied to clipboard
Abdellah EL Mekki, Samar M. Magdy, Houdaifa Atou, Ruwa AbuHweidi, Baraah Qawasmeh, Omer Nacar, Thikra Al-hibiri, Razan Saadie, Hamzah A. Alsayadi, Nadia Ghezaiel Hammouda, Alshima Mohammed Alkhazimi, Aya Hamod, Al-Yas Yaqoob Al-Ghafri, Wesam El-Sayed, Asila Ismail al Sharji, Mohamad Ballout, Anas Belfathi, Karim Ghaddar, Serry Sibaee, Alaa Aoun, Aeej Mohammed Aseri, Lina Abureesh, Ahlam Bashiti, Majdal Yousef, Abdulaziz Hafiz, Yehdih Mohamed, Emira Hamedtou, Brakehe Emehah, Rahaf Alhamouri, Youssef Nafea, Aya El Aatar, Walid Al-Dhabyani, Emhemed S. Hamed, Sara Shatnawi, Fakhraddin Alwajih, Khalid Elkhidir, Ashwag Alasmari, Abdurrahman Gerrio, Omar Said Alshahri, AbdelRahim A. Elmadany, Ismail Berrada, Amir Azad Adli Al-kathiri, Fadi Zaraket, Mustafa Jarrar, Yahya Mohamed EL Hadj, Hassan Alhuzali, Muhammad Abdul-Mageed
| Challenge: | Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than modern standard Arabic (MSA). |
| Approach: | They propose a large-scale, community-driven, human-translated dataset to bridge this gap . Alexandria covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . |
| Outcome: | The Alexandria dataset covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . Alexandria is a training resource and a rigorous benchmark for evaluating MT and LLMs based on the Alexandria dataset . |
NeoAraBERT: A Modern Foundation Model for Arabic Embeddings with Diacritics-Aware Tokenization and POS-Targeted Masking (2026.findings-acl)
Copied to clipboard
Chadi Abou Chakra, Hadi Khaled Hamoud, Osama Rakan Al Mraikhat, Qusai Abu Obaida, Mohamad Ballout, Fadi Zaraket
| Challenge: | NeoAraBERT is an open-source text-embedding model for Arabic. |
| Approach: | They propose to train Arabic text-embedding models on open-source datasets . they benchmarked NeoAraBERT against five top-performing Arabic models on 23 tasks . |
| Outcome: | The proposed model outperforms five other models on 23 tasks in Arabic . it shows substantial improvement on classical and modern standard Arabic compared to other models . |
DAVE: Differential Diagnostic Analysis Automation and Visualization from Clinical Notes (2023.eacl-demo)
Copied to clipboard
| Challenge: | a tool that uses natural language processing and machine learning to help visualize diagnostic algorithms in real-time. |
| Approach: | They propose a tool that uses natural language processing and machine learning to visualize diagnostic algorithms in real-time. |
| Outcome: | The proposed system automates the selection and visualization process of diagnostic algorithms. |
Curras + Baladi: Towards a Levantine Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | The processing of the Arabic language is a complex field of research due to the complex and rich morphology of Arabic, its high degree of ambiguity, and the presence of several regional varieties that need to be processed while taking into account their unique characteristics. |
| Approach: | They propose to revise the Palestinian morphologically annotated corpus and a Lebanese corpus to bridge nuanced linguistic gaps between the two highly mutually intelligible dialects. |
| Outcome: | The revised corpus can be used as a more general Levantine corpus. |
R-BPE: Improving BPE-Tokenizers with Token Reuse (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large pretrained language models prioritize high-resource languages in their vocabularies, leaving others with poor coverage. |
| Approach: | They propose a framework that reuses existing tokenizers and creates ID-based maps to resolve the new tokens of the chosen language. |
| Outcome: | The proposed framework reduces subword fertility by 24.4% on Arabic models and preserves performance on EnglishMMLU. |