Challenge: a novel image-based character embedding framework is used for text classification in Arabic . classical methods require morphological analysis, word segmentation, and hand-crafted feature engineering.
Approach: They propose a novel end-to-end Arabic document classification framework, Arabic document image-based classifier, inspired by image-basic character embeddings.
Outcome: The proposed framework improves on modern standard Arabic, colloquial Arabic, and Classical Arabic.

Similar Papers

KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Optical Character Recognition (OCR) is a key component of document processing . Arabic text recognition has complex typographic and calligraphic features .
Approach: They propose a comprehensive Arabic OCR benchmark that fills the gaps in evaluation systems.
Outcome: The proposed benchmark outperforms existing models in Arabic by 60% in the character error rate . the best model achieves only 65% accuracy in PDF-to-Markdown conversion .
AraMUS: Pushing the Limits of Data and Model Scale for Arabic Natural Language Processing (2023.findings-acl)

Copied to clipboard

Challenge: Developing monolingual large Pre-trained Language Models (PLMs) is shown to be very successful in handling different tasks in Natural Language Processing (NLP).
Approach: They present AraMUS, the largest Arabic PLM with 11B parameters trained on 529GB of high-quality Arabic textual data.
Outcome: The proposed model achieves state-of-the-art performance on a diverse set of Arabic classification and generative tasks.
AraT5: Text-to-Text Transformers for Arabic Language Generation (2022.acl-long)

Copied to clipboard

Challenge: Existing models that convert text-based language problems into text-to-text format are not suitable for multilingual tasks.
Approach: They propose a unified Transformer framework that converts all language problems into a text-to-text format.
Outcome: The proposed model performs better on all ARGEN tasks than existing models with 49 less data.
Enhancing Deep Learning with Embedded Features for Arabic Named Entity Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Word embeddings can capture the semantics of words and other hidden features, but the Arabic language is complex and requires a large amount of information to process.
Approach: They propose to add morphological and syntactical features to Arabic word embeddings to train the model.
Outcome: The proposed model outperforms the previous systems to the best of our knowledge.
AdabNER: Arabic Digital Archive Books with Nested Entity Recognition (2026.acl-long)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a subtask of information extraction that classifies entities into predefined categories like person names.
Approach: They propose a large-scale nested Arabic Named Entity Recognition dataset . they fine-tuned five pre-trained Arabic BERT encoders in two settings .
Outcome: The first large-scale nested NER dataset for Arabic literary texts is published online . the dataset yields 78,530 entity mentions, 18.96% of which are nestated .
DAICT: A Dialectal Arabic Irony Corpus Extracted from Twitter (2020.lrec-1)

Copied to clipboard

Challenge: Current scholarship is yet to reach an agreement on a universal definition of the concept of irony.
Approach: They propose to query Twitter using irony-related hashtags to collect ironic messages which are then manually annotated by two linguists according to their working definition of irony.
Outcome: The proposed corpus will be a valuable resource for developing open domain systems for automatic irony recognition in Arabic and its dialects in social media text.
Enhancing Arabic NLP Tasks through Character-Level Models and Data Augmentation (2025.coling-main)

Copied to clipboard

Challenge: Using character-level models, natural language processing for Arabic is challenging due to its rich morphology, root-based word formation, flexible sentence structures, diacritical ambiguities, and orthographic variations.
Approach: They propose a character-level approach specifically designed for Arabic NLP tasks that incorporates Convolutional Neural Networks (CNNs), pre-trained transformers (CANINE), and Bidirectional Long Short-Term Memory networks (BiLSTMs).
Outcome: The proposed model outperforms existing models on Arabic privacy policy classification task and reports a micro-averaged F1 score of 93.8%, surpassing state-of-the-art models.
Toward Qualitative Evaluation of Embeddings for Arabic Sentiment Analysis (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies on Arabic sentiment analysis (SA) tasks focus on word embeddings to capture semantic and syntactic similarities, but Arabic language is characterized by its agglutination and morphological richness contributing to great sparsity.
Approach: They propose several protocols to evaluate specific embeddings for Arabic sentiment analysis task.
Outcome: The proposed embeddings are based on words and lemmas in Arabic sentiment analysis (SA) task.
Normalising Non-standardised Orthography in Algerian Code-switched User-generated Data (D19-55)

Copied to clipboard

Challenge: a new corpus of unstructured data from social media is presenting challenges to NLP research . standardisation is neither natural nor universal, it is rather a human invention.
Approach: They compile a parallel corpus of Arabic textual data matched with human annotations . they use a deep neural model designed to deal with context-dependent spelling correction .
Outcome: The proposed model performs best with two CNN sub-network encoders and an LSTM decoder . pre-processing data token-by-token with edit-distance aligner significantly improves performance .
Neural Arabic Text Diacritization: State of the Art Results and a Novel Approach for Machine Translation (D19-52)

Copied to clipboard

Challenge: a number of Arabic text diacritizers use diacritics to convey information about meaning of a word . Arabic text to speech (TTS) requires a complex process to determine the correct diacritical for each character .
Approach: They propose to use Arabic diacritization to enhance machine translation models . they propose to build automatic Arabic text diacritics using two approaches .
Outcome: The proposed models are either better or on par with other models, which require language-dependent post-processing steps, unlike ours.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations