Challenge: Existing evaluation datasets feature Western-centric images and English text, while their non-English counterparts are often derived from the latter.
Approach: They propose to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Morocco.
Outcome: The proposed model underperforms in visual understanding and dialect-specific generation across four Arabic-speaking countries.

Similar Papers

Benchmarking Vision Language Models for Cultural Understanding (2024.emnlp-main)

Copied to clipboard

Challenge: Recent multimodal vision-language models have shown impressive performance in tasks such as image-to-text generation, visual question answering, and image captioning.
Approach: They propose a visual question-answering benchmark to assess VLMs' cultural understanding of various facets of culture from 11 countries across 5 continents.
Outcome: The visual question-answering benchmark aims to assess VLMs' cultural understanding across regions.
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs (2025.coling-main)

Copied to clipboard

Challenge: a recent study has found that Arabic is underrepresented in Large Language Models, especially in dialectal variations.
Approach: They propose a benchmark for Arabic Dialect and Cultural Evaluation that evaluates Arabic dialect comprehension and generation.
Outcome: The proposed model outperforms multilingual models on dialect comprehension and generation, but significant challenges persist in dialect identification, generation, and translation.
AL-QASIDA: Analyzing LLM Quality and Accuracy Systematically in Dialectal Arabic (2025.findings-acl)

Copied to clipboard

Challenge: Dialectal Arabic (DA) varieties are under-served by language technologies, particularly large language models (LLMs).
Approach: They propose a framework that comprehensively assesses LLMs’ DA modeling capabilities across four dimensions: fidelity, understanding, quality, and diglossia.
Outcome: The proposed framework assesses LLMs’ DA modeling capabilities across four dimensions: fidelity, understanding, quality, and diglossia.
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset (2025.findings-emnlp)

Copied to clipboard

Challenge: Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets.
Approach: They propose to construct a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding.
Outcome: The proposed dataset covers ten culturally significant domains covering all Arab countries and includes two evaluation benchmarks (PEARL and PEARL-LITE) and a specialized subset (PearL-X).
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark (2025.findings-naacl)

Copied to clipboard

Challenge: Recent years have witnessed a significant interest in developing large multimodal models capable of performing various visual reasoning and understanding tasks.
Approach: They propose to use Arabic as a language to evaluate large multi-modal models capable of performing visual reasoning and understanding tasks.
Outcome: The proposed benchmark comprises eight diverse domains and 38 sub-domains to represent a large population of over 400 million speakers.
JAWAHER: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking (2025.naacl-long)

Copied to clipboard

Challenge: Recent advances in instruction fine-tuning and alignment methods have enhanced the adaptability of large language models to user preferences.
Approach: They propose a benchmark to assess LLMs’ capacity to comprehend and interpret Arabic proverbs.
Outcome: The proposed model can generate accurate translations, but struggle to produce culturally nuanced and contextually relevant explanations.
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic (2024.findings-acl)

Copied to clipboard

Challenge: evaluating language models in Arabic remains challenging due to limited datasets . focus has shift to reasoning and knowledge-intensive tasks due to lack of relevant datasets.
Approach: They propose to use ArabicMMLU to evaluate models' understanding of Arabic . they use 40 tasks and 14,575 multiple-choice questions from school exams in different countries .
Outcome: The ArabicMMLU is the first multi-task language understanding benchmark for the Arabic language . it is based on 40 tasks and 14,575 multiple-choice questions in modern standard Arabic . the models are based in different countries across North Africa, the Levant, and the Gulf regions .
Nahw: A Comprehensive Benchmark of Arabic Grammar Understanding, Error Detection, Correction, and Explanation (2026.eacl-long)

Copied to clipboard

Challenge: Existing corpora address individual linguistic aspects like spelling or diacritization, but rarely provide explanations of grammatical errors. Existing datasets and benchmarks that capture Arabic's grammatological complexity are scarce.
Approach: They propose a benchmark for Arabic grammar that covers error detection, correction, and explanation.
Outcome: The proposed model performs better on GPT-4o than on the best performing model (ALLaM-7B) despite fine tuning with synthetic data, the model perform better on Arabic grammar tasks.
Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks (2025.findings-naacl)

Copied to clipboard

Challenge: In this paper, we introduce a family of embedding models addressing both small-scale and large-scale use cases.
Approach: They propose to use ArabicMTEB to evaluate Arabic text embedding models . they propose to build a benchmark suite that assesses cross-lingual, multi-dialectal, multidomain, and multi-cultural Arabic text embedded models.
Outcome: The proposed models outperform Multilingual-E5-large and Swan-Large in most Arabic tasks while remaining dialectally and culturally aware.
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology (2026.findings-acl)

Copied to clipboard

Challenge: Dialectal Arabic datasets embody a range of domain, dialect, and quality.
Approach: They propose a framework for automatic speech recognition in dialectal Arabic to address the limited data availability encountered in dialects.
Outcome: The proposed framework provides access to 31 datasets covering 14 dialects to better address the limited data availability encountered in dialectal Arabic speech processing.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations