Sara Ghaboura, Ahmed Heakl, Omkar Thawakar, Ali Husain Salem Abdulla Alharthi, Ines Riahi, Abduljalil Radman, Jorma Laaksonen, Fahad Shahbaz Khan, Salman Khan, Rao Muhammad Anwer
| Challenge: | Recent years have witnessed a significant interest in developing large multimodal models capable of performing various visual reasoning and understanding tasks. |
| Approach: | They propose to use Arabic as a language to evaluate large multi-modal models capable of performing visual reasoning and understanding tasks. |
| Outcome: | The proposed benchmark comprises eight diverse domains and 38 sub-domains to represent a large population of over 400 million speakers. |
Similar Papers
KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding (2025.findings-acl)
Copied to clipboard
Ahmed Heakl, Muhammad Abdullah Sohail, Mukul Ranjan, Rania Elbadry, Ghazi Shazan Ahmad, Mohamed El-Geish, Omar Maher, Zhiqiang Shen, Fahad Shahbaz Khan, Salman Khan
| Challenge: | Optical Character Recognition (OCR) is a key component of document processing . Arabic text recognition has complex typographic and calligraphic features . |
| Approach: | They propose a comprehensive Arabic OCR benchmark that fills the gaps in evaluation systems. |
| Outcome: | The proposed benchmark outperforms existing models in Arabic by 60% in the character error rate . the best model achieves only 65% accuracy in PDF-to-Markdown conversion . |
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic (2024.findings-acl)
Copied to clipboard
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, Timothy Baldwin
| Challenge: | evaluating language models in Arabic remains challenging due to limited datasets . focus has shift to reasoning and knowledge-intensive tasks due to lack of relevant datasets. |
| Approach: | They propose to use ArabicMMLU to evaluate models' understanding of Arabic . they use 40 tasks and 14,575 multiple-choice questions from school exams in different countries . |
| Outcome: | The ArabicMMLU is the first multi-task language understanding benchmark for the Arabic language . it is based on 40 tasks and 14,575 multiple-choice questions in modern standard Arabic . the models are based in different countries across North Africa, the Levant, and the Gulf regions . |
TounsiBench: Benchmarking Large Language Models for Tunisian Arabic (2025.emnlp-main)
Copied to clipboard
| Challenge: | a dataset of Tunisian Arabic instructions and prompts is used to evaluate LLMs' ability to understand and generate responses in Tunisia . we assess the quality, correctness, relevance, and dialectal adherence of LLM responses . |
| Approach: | They propose a benchmark for evaluating the capabilities of large language models in Tunisian Arabic . they use a dataset of Tunisia Arabic instructions and prompts to evaluate their models . |
| Outcome: | The proposed model can judge quality, correctness, relevance, and dialectal adherence . the model can also generate a leaderboard for the Tunisian Arabic language . |
Nahw: A Comprehensive Benchmark of Arabic Grammar Understanding, Error Detection, Correction, and Explanation (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing corpora address individual linguistic aspects like spelling or diacritization, but rarely provide explanations of grammatical errors. Existing datasets and benchmarks that capture Arabic's grammatological complexity are scarce. |
| Approach: | They propose a benchmark for Arabic grammar that covers error detection, correction, and explanation. |
| Outcome: | The proposed model performs better on GPT-4o than on the best performing model (ALLaM-7B) despite fine tuning with synthetic data, the model perform better on Arabic grammar tasks. |
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset (2025.findings-emnlp)
Copied to clipboard
Fakhraddin Alwajih, Samar M. Magdy, Abdellah El Mekki, Omer Nacar, Youssef Nafea, Safaa Taher Abdelfadil, Abdulfattah Mohammed Yahya, Hamzah Luqman, Nada Almarwani, Samah Aloufi, Baraah Qawasmeh, Houdaifa Atou, Serry Sibaee, Hamzah A. Alsayadi, Walid Al-Dhabyani, Maged S. Al-shaibani, Aya El aatar, Nour Qandos, Rahaf Alhamouri, Samar Ahmad, Mohammed Anwar AL-Ghrawi, Aminetou Yacoub, Ruwa AbuHweidi, Vatimetou Mohamed Lemin, Reem Abdel-Salam, Ahlam Bashiti, Adel Ammar, Aisha Alansari, Ahmed Ashraf, Nora Alturayeif, Alcides Alcoba Inciarte, AbdelRahim A. Elmadany, Mohamedou Cheikh Tourad, Ismail Berrada, Mustafa Jarrar, Shady Shehata, Muhammad Abdul-Mageed
| Challenge: | Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. |
| Approach: | They propose to construct a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding. |
| Outcome: | The proposed dataset covers ten culturally significant domains covering all Arab countries and includes two evaluation benchmarks (PEARL and PEARL-LITE) and a specialized subset (PearL-X). |
LAraBench: Benchmarking Arabic AI with Large Language Models (2024.eacl-long)
Copied to clipboard
Ahmed Abdelali, Hamdy Mubarak, Shammur Chowdhury, Maram Hasanain, Basel Mousi, Sabri Boughorbel, Samir Abdaljalil, Yassine El Kheir, Daniel Izham, Fahim Dalvi, Majd Hawasly, Nizi Nazar, Youssef Elshahawy, Ahmed Ali, Nadir Durrani, Natasa Milic-Frayling, Firoj Alam
| Challenge: | Recent advances in Large Language Models (LLMs) have significantly influenced the landscape of language and speech research. |
| Approach: | They used GPT-3.5-turbo, GPT-4, BLOOMZ, Jais-13b-chat, Whisper, and USM to tackle 33 distinct tasks across 61 datasets. |
| Outcome: | The proposed model outperforms SOTA models in zero-shot learning, with a few exceptions. |
AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic (2025.coling-main)
Copied to clipboard
| Challenge: | Existing benchmarks for large language models (LLMs) in Arabic are lacking . despite progress in their development, there is a lack of comprehensive trustworthiness evaluation benchmarks . |
| Approach: | They propose to use Arabic as a language to assess trustworthiness of large language models. |
| Outcome: | The proposed benchmark measures the trustworthiness of large language models in Arabic. |
ORCA: A Challenging Benchmark for Arabic Language Understanding (2023.findings-acl)
Copied to clipboard
| Challenge: | Despite efforts to evaluate Arabic NLU, no public benchmark of diverse nature exists . a benchmark targeting Arabic needs to take into account that Arabic is not a single language but a collection of languages and language varieties. |
| Approach: | They propose a publicly available benchmark for Arabic language understanding evaluation dubbed ORCA . it covers diverse Arabic varieties and a wide range of Arabic understanding tasks . |
| Outcome: | The proposed benchmark covers Arabic and multilingual models across seven NLU task clusters. |
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models (2025.findings-naacl)
Copied to clipboard
Kaichen Zhang, Bo Li, Peiyuan Zhang, Fanyi Pu, Joshua Adrian Cahyono, Kairui Hu, Shuai Liu, Yuanhan Zhang, Jingkang Yang, Chunyuan Li, Ziwei Liu
| Challenge: | Current large foundational models have demonstrated transformative capabilities, approaching or surpassing human-level performances in many tasks. |
| Approach: | They propose a unified and standardized multimodal benchmark framework with over 50 tasks and more than 10 models to promote transparent and reproducible evaluations. |
| Outcome: | The proposed framework has 50 tasks and more than 10 models to promote transparent and reproducible evaluations. |
JEEM: Vision-Language Understanding in Four Arabic Dialects (2026.findings-eacl)
Copied to clipboard
Karima Kadaoui, Hanin Atwany, Hamdan Al-Ali, Abdelrahman Mohamed, Ali Mekky, Sergei Tilga, Natalia Fedorova, Ekaterina Artemova, Hanan Aldarmaki, Yova Kementchedjhieva
| Challenge: | Existing evaluation datasets feature Western-centric images and English text, while their non-English counterparts are often derived from the latter. |
| Approach: | They propose to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Morocco. |
| Outcome: | The proposed model underperforms in visual understanding and dialect-specific generation across four Arabic-speaking countries. |