Challenge: Optical Character Recognition (OCR) is a key component of document processing . Arabic text recognition has complex typographic and calligraphic features .
Approach: They propose a comprehensive Arabic OCR benchmark that fills the gaps in evaluation systems.
Outcome: The proposed benchmark outperforms existing models in Arabic by 60% in the character error rate . the best model achieves only 65% accuracy in PDF-to-Markdown conversion .

Similar Papers

CAMEL-Bench: A Comprehensive Arabic LMM Benchmark (2025.findings-naacl)

Copied to clipboard

Challenge: Recent years have witnessed a significant interest in developing large multimodal models capable of performing various visual reasoning and understanding tasks.
Approach: They propose to use Arabic as a language to evaluate large multi-modal models capable of performing visual reasoning and understanding tasks.
Outcome: The proposed benchmark comprises eight diverse domains and 38 sub-domains to represent a large population of over 400 million speakers.
Nahw: A Comprehensive Benchmark of Arabic Grammar Understanding, Error Detection, Correction, and Explanation (2026.eacl-long)

Copied to clipboard

Challenge: Existing corpora address individual linguistic aspects like spelling or diacritization, but rarely provide explanations of grammatical errors. Existing datasets and benchmarks that capture Arabic's grammatological complexity are scarce.
Approach: They propose a benchmark for Arabic grammar that covers error detection, correction, and explanation.
Outcome: The proposed model performs better on GPT-4o than on the best performing model (ALLaM-7B) despite fine tuning with synthetic data, the model perform better on Arabic grammar tasks.
Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks (2025.findings-naacl)

Copied to clipboard

Challenge: In this paper, we introduce a family of embedding models addressing both small-scale and large-scale use cases.
Approach: They propose to use ArabicMTEB to evaluate Arabic text embedding models . they propose to build a benchmark suite that assesses cross-lingual, multi-dialectal, multidomain, and multi-cultural Arabic text embedded models.
Outcome: The proposed models outperform Multilingual-E5-large and Swan-Large in most Arabic tasks while remaining dialectally and culturally aware.
Advancing Arabic Diacritization: Improved Datasets, Benchmarking, and State-of-the-Art Models (2025.emnlp-main)

Copied to clipboard

Challenge: Arabic diacritics are typically omitted in written Arabic, leading to ambiguity . authors propose a methodology to analyze and refine a large diacritized corpus .
Approach: They propose a methodology to analyze and refine a large diacritized corpus to improve training quality.
Outcome: The proposed model achieves state-of-the-art results with 3.12% and 2.70% WER on WikiNews-2014 and Wikinews-2024.
ORCA: A Challenging Benchmark for Arabic Language Understanding (2023.findings-acl)

Copied to clipboard

Challenge: Despite efforts to evaluate Arabic NLU, no public benchmark of diverse nature exists . a benchmark targeting Arabic needs to take into account that Arabic is not a single language but a collection of languages and language varieties.
Approach: They propose a publicly available benchmark for Arabic language understanding evaluation dubbed ORCA . it covers diverse Arabic varieties and a wide range of Arabic understanding tasks .
Outcome: The proposed benchmark covers Arabic and multilingual models across seven NLU task clusters.
DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding (2026.eacl-long)

Copied to clipboard

Challenge: a benchmark of 1,272 samples containing about 1,475 unique words is available for Arabic calligraphy . the dataset reflects real-world challenges in Arabic writing, such as calligraphic variation and artistic distortions .
Approach: They evaluated 13 leading Arabic and multilingual multimodal models and paired them with sentence-level annotations to evaluate their calligraphy models.
Outcome: The benchmark evaluates 13 leading Arabic and multilingual multimodal models . it shows they struggle with calligraphic variation, artistic distortions, and precise visual–text alignment.
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset (2025.findings-emnlp)

Copied to clipboard

Challenge: Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets.
Approach: They propose to construct a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding.
Outcome: The proposed dataset covers ten culturally significant domains covering all Arab countries and includes two evaluation benchmarks (PEARL and PEARL-LITE) and a specialized subset (PearL-X).
Tafsir Dataset: A Novel Multi-Task Benchmark for Named Entity Recognition and Topic Modeling in Classical Arabic Literature (2022.coling-1)

Copied to clipboard

Challenge: Named entity recognition and topic modeling are crucial for downstream tasks in natural language processing.
Approach: They propose to address named entity recognition and topic modeling on CA literature . they manually annotate the work of Tafsir Al-Tabari with span-based NEs .
Outcome: The results show that the proposed task can perform state-of-the-art on historical topic models.
TounsiBench: Benchmarking Large Language Models for Tunisian Arabic (2025.emnlp-main)

Copied to clipboard

Challenge: a dataset of Tunisian Arabic instructions and prompts is used to evaluate LLMs' ability to understand and generate responses in Tunisia . we assess the quality, correctness, relevance, and dialectal adherence of LLM responses .
Approach: They propose a benchmark for evaluating the capabilities of large language models in Tunisian Arabic . they use a dataset of Tunisia Arabic instructions and prompts to evaluate their models .
Outcome: The proposed model can judge quality, correctness, relevance, and dialectal adherence . the model can also generate a leaderboard for the Tunisian Arabic language .
Konooz: Multi-domain Multi-dialect Corpus for Named Entity Recognition (2025.findings-acl)

Copied to clipboard

Challenge: Using the Wojood framework, we compare existing Arabic Named Entity Recognition models with domain and dialect divergence and resource scarcity.
Approach: They propose a multi-dimensional Arabic named entity corpus covering 16 dialects across 10 domains and an annotation scheme using the Wojood guidelines.
Outcome: The proposed model performs better on 16 dialects across 10 domains and 16 domains, while other models struggle with different dialects and domains.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations