Papers by Mohamad Ballout

5 papers
Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs (2026.acl-long)

Copied to clipboard

Challenge: Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than modern standard Arabic (MSA).
Approach: They propose a large-scale, community-driven, human-translated dataset to bridge this gap . Alexandria covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata .
Outcome: The Alexandria dataset covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . Alexandria is a training resource and a rigorous benchmark for evaluating MT and LLMs based on the Alexandria dataset .
NeoAraBERT: A Modern Foundation Model for Arabic Embeddings with Diacritics-Aware Tokenization and POS-Targeted Masking (2026.findings-acl)

Copied to clipboard

Challenge: NeoAraBERT is an open-source text-embedding model for Arabic.
Approach: They propose to train Arabic text-embedding models on open-source datasets . they benchmarked NeoAraBERT against five top-performing Arabic models on 23 tasks .
Outcome: The proposed model outperforms five other models on 23 tasks in Arabic . it shows substantial improvement on classical and modern standard Arabic compared to other models .
Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: SPLICE is a benchmark designed to probe event-based reasoning across multiple dimensions.
Approach: They introduce a human-curated benchmark to probe event-based reasoning across multiple dimensions.
Outcome: The proposed benchmark includes 3,381 human-filtered videos spanning 12 categories and 180 sub-categories . results show that state-of-the-art vision-language models struggle to match human performance .
FOOL ME IF YOU CAN! An Adversarial Dataset to Investigate the Robustness of LMs in Word Sense Disambiguation (2024.emnlp-main)

Copied to clipboard

Challenge: Word sense disambiguation (WSD) is a key task in natural language processing . however, these models struggle with recognizing semantic boundaries in adversarial contexts .
Approach: They propose to use a coarse-grained WSD dataset to assess model robustness . they found that some models struggled to correctly disambiguate homonyms in adversarial contexts .
Outcome: The proposed dataset includes four test sets to assess the robustness of language models in WSD tasks.
iVISPAR — An Interactive Visual-Spatial Reasoning Benchmark for VLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Vision-Language Models (VLMs) struggle with spatial reasoning and visual alignment, despite their performance on 2D tasks.
Approach: They propose a multimodal benchmark to evaluate VLMs' spatial reasoning capabilities based on the sliding tile puzzle .
Outcome: The proposed model performs better on 2D tasks compared to 3D or text-based settings, but struggles with complex spatial configurations and consistently falls short of human performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations