ORCA: A Challenging Benchmark for Arabic Language Understanding (2023.findings-acl)
Copied to clipboard
| Challenge: | Despite efforts to evaluate Arabic NLU, no public benchmark of diverse nature exists . a benchmark targeting Arabic needs to take into account that Arabic is not a single language but a collection of languages and language varieties. |
| Approach: | They propose a publicly available benchmark for Arabic language understanding evaluation dubbed ORCA . it covers diverse Arabic varieties and a wide range of Arabic understanding tasks . |
| Outcome: | The proposed benchmark covers Arabic and multilingual models across seven NLU task clusters. |
Similar Papers
Dolphin: A Challenging and Diverse Benchmark for Arabic NLG (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing benchmarks for Arabic are limited, but they can be used to measure performance of different languages. |
| Approach: | They propose a benchmark for Arabic that addresses the need for a framework dedicated to Arabic languages and varieties. |
| Outcome: | The proposed benchmark covers 13 different tasks in Arabic and spans 50 test splits. |
Nahw: A Comprehensive Benchmark of Arabic Grammar Understanding, Error Detection, Correction, and Explanation (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing corpora address individual linguistic aspects like spelling or diacritization, but rarely provide explanations of grammatical errors. Existing datasets and benchmarks that capture Arabic's grammatological complexity are scarce. |
| Approach: | They propose a benchmark for Arabic grammar that covers error detection, correction, and explanation. |
| Outcome: | The proposed model performs better on GPT-4o than on the best performing model (ALLaM-7B) despite fine tuning with synthetic data, the model perform better on Arabic grammar tasks. |
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic (2024.findings-acl)
Copied to clipboard
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, Timothy Baldwin
| Challenge: | evaluating language models in Arabic remains challenging due to limited datasets . focus has shift to reasoning and knowledge-intensive tasks due to lack of relevant datasets. |
| Approach: | They propose to use ArabicMMLU to evaluate models' understanding of Arabic . they use 40 tasks and 14,575 multiple-choice questions from school exams in different countries . |
| Outcome: | The ArabicMMLU is the first multi-task language understanding benchmark for the Arabic language . it is based on 40 tasks and 14,575 multiple-choice questions in modern standard Arabic . the models are based in different countries across North Africa, the Levant, and the Gulf regions . |
TounsiBench: Benchmarking Large Language Models for Tunisian Arabic (2025.emnlp-main)
Copied to clipboard
| Challenge: | a dataset of Tunisian Arabic instructions and prompts is used to evaluate LLMs' ability to understand and generate responses in Tunisia . we assess the quality, correctness, relevance, and dialectal adherence of LLM responses . |
| Approach: | They propose a benchmark for evaluating the capabilities of large language models in Tunisian Arabic . they use a dataset of Tunisia Arabic instructions and prompts to evaluate their models . |
| Outcome: | The proposed model can judge quality, correctness, relevance, and dialectal adherence . the model can also generate a leaderboard for the Tunisian Arabic language . |
Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language Processing (2022.emnlp-main)
Copied to clipboard
Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Chao Xing, Yasheng Wang, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Xin Jiang, Qun Liu, Phillippe Langlais
| Challenge: | Existing pre-trained language models are not well-explored and are not reproducible in the literature. |
| Approach: | They propose to improve existing Arabic language pre-trained language models using a more methodical approach. |
| Outcome: | The proposed models outperform existing models on ALUE, a leaderboard-powered benchmark for Arabic NLU and NLG tasks. |
Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs (2025.emnlp-main)
Copied to clipboard
Wafa Al Ghallabi, Ritesh Thawkar, Sara Ghaboura, Ketan Pravin More, Omkar Thawakar, Hisham Cholakkal, Salman Khan, Rao Muhammad Anwer
| Challenge: | a benchmark is designed to assess the comprehension of Arabic poetry by large language models in 12 historical eras. |
| Approach: | They propose a benchmark to assess the comprehension of Arabic poetry by large language models in 12 historical eras. |
| Outcome: | The benchmark assesses the comprehension of Arabic poetry by large language models in 12 historical eras. |
LAraBench: Benchmarking Arabic AI with Large Language Models (2024.eacl-long)
Copied to clipboard
Ahmed Abdelali, Hamdy Mubarak, Shammur Chowdhury, Maram Hasanain, Basel Mousi, Sabri Boughorbel, Samir Abdaljalil, Yassine El Kheir, Daniel Izham, Fahim Dalvi, Majd Hawasly, Nizi Nazar, Youssef Elshahawy, Ahmed Ali, Nadir Durrani, Natasa Milic-Frayling, Firoj Alam
| Challenge: | Recent advances in Large Language Models (LLMs) have significantly influenced the landscape of language and speech research. |
| Approach: | They used GPT-3.5-turbo, GPT-4, BLOOMZ, Jais-13b-chat, Whisper, and USM to tackle 33 distinct tasks across 61 datasets. |
| Outcome: | The proposed model outperforms SOTA models in zero-shot learning, with a few exceptions. |
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing multilingual benchmarks focus primarily on language understanding tasks. |
| Approach: | They develop a multi-way multilingual benchmark that measures critical capabilities of large language models across languages. |
| Outcome: | Extensive experiments on BenchMAX reveal uneven utilization of core capabilities across languages, emphasizing the performance gaps that scaling model size alone does not resolve. |
The SADID Evaluation Datasets for Low-Resource Spoken Language Machine Translation of Arabic Dialects (2020.coling-main)
Copied to clipboard
| Challenge: | Low-resource Machine Translation (LRT) models are still lagging behind on low-resourced language pairs due to the scarcity of parallel training data. |
| Approach: | They introduce benchmark datasets for Arabic and its dialects to examine their properties . they bootstrap existing parallel sentences and complement this with multilingual training . |
| Outcome: | The proposed method bootstraps existing parallel sentences and complements multilingual training to achieve strong baselines. |
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark (2025.findings-naacl)
Copied to clipboard
Sara Ghaboura, Ahmed Heakl, Omkar Thawakar, Ali Husain Salem Abdulla Alharthi, Ines Riahi, Abduljalil Radman, Jorma Laaksonen, Fahad Shahbaz Khan, Salman Khan, Rao Muhammad Anwer
| Challenge: | Recent years have witnessed a significant interest in developing large multimodal models capable of performing various visual reasoning and understanding tasks. |
| Approach: | They propose to use Arabic as a language to evaluate large multi-modal models capable of performing visual reasoning and understanding tasks. |
| Outcome: | The proposed benchmark comprises eight diverse domains and 38 sub-domains to represent a large population of over 400 million speakers. |