Papers by Fadi Zaraket

6 papers
From RAG to Agentic RAG for Faithful Islamic Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used for Islamic question answering, where ungrounded responses may carry serious religious consequences.
Approach: They propose a bilingual, bilingual, Arabic/English benchmark with atomic single-gold answers that measures hallucination and abstention.
Outcome: The proposed model improves accuracy and robustness even with a small model.
Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs (2026.acl-long)

Copied to clipboard

Challenge: Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than modern standard Arabic (MSA).
Approach: They propose a large-scale, community-driven, human-translated dataset to bridge this gap . Alexandria covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata .
Outcome: The Alexandria dataset covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . Alexandria is a training resource and a rigorous benchmark for evaluating MT and LLMs based on the Alexandria dataset .
NeoAraBERT: A Modern Foundation Model for Arabic Embeddings with Diacritics-Aware Tokenization and POS-Targeted Masking (2026.findings-acl)

Copied to clipboard

Challenge: NeoAraBERT is an open-source text-embedding model for Arabic.
Approach: They propose to train Arabic text-embedding models on open-source datasets . they benchmarked NeoAraBERT against five top-performing Arabic models on 23 tasks .
Outcome: The proposed model outperforms five other models on 23 tasks in Arabic . it shows substantial improvement on classical and modern standard Arabic compared to other models .
DAVE: Differential Diagnostic Analysis Automation and Visualization from Clinical Notes (2023.eacl-demo)

Copied to clipboard

Challenge: a tool that uses natural language processing and machine learning to help visualize diagnostic algorithms in real-time.
Approach: They propose a tool that uses natural language processing and machine learning to visualize diagnostic algorithms in real-time.
Outcome: The proposed system automates the selection and visualization process of diagnostic algorithms.
Curras + Baladi: Towards a Levantine Corpus (2022.lrec-1)

Copied to clipboard

Challenge: The processing of the Arabic language is a complex field of research due to the complex and rich morphology of Arabic, its high degree of ambiguity, and the presence of several regional varieties that need to be processed while taking into account their unique characteristics.
Approach: They propose to revise the Palestinian morphologically annotated corpus and a Lebanese corpus to bridge nuanced linguistic gaps between the two highly mutually intelligible dialects.
Outcome: The revised corpus can be used as a more general Levantine corpus.
R-BPE: Improving BPE-Tokenizers with Token Reuse (2025.emnlp-main)

Copied to clipboard

Challenge: Large pretrained language models prioritize high-resource languages in their vocabularies, leaving others with poor coverage.
Approach: They propose a framework that reuses existing tokenizers and creates ID-based maps to resolve the new tokens of the chosen language.
Outcome: The proposed framework reduces subword fertility by 24.4% on Arabic models and preserves performance on EnglishMMLU.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations