Challenge: Using the internet, the spoken Arabic dialect language becomes informal languages written in social media . this linguistic situation inhibits mutual understanding and makes computational approaches difficult . we present a pipeline to standardize the written texts in social networks by translating them to MSA .
Approach: They propose a pipeline to standardize Arabic written texts by translating them to MSA . they use a bert-based model to select Tunisian Dialect from MSA and other dialects .
Outcome: The proposed pipeline achieves the best score for the standardization of written texts in social networks . the proposed pipeline includes the translated TD and the original text written in MSA .

Similar Papers

ALDi: Quantifying the Arabic Level of Dialectness of Text (2023.emnlp-main)

Copied to clipboard

Challenge: Existing work on Dialect Identification (DI) on the sentence level has focused on binary tasks, whereas ALDi treats the task as binary.
Approach: They propose a dataset which contains 127,835 sentences manually labeled with their level of dialectness.
Outcome: The proposed model can identify dialectness on a range of other corpora, providing a more nuanced picture than traditional DI systems.
Shami: A Corpus of Levantine Arabic Dialects (L18-1)

Copied to clipboard

Challenge: Modern Standard Arabic is the official written language used in education and media . however, the spoken language varies widely across the Arab world .
Approach: They construct a levantine dialect corpus covering data from four dialects spoken in four countries . they describe rules for pre-processing without affecting the meaning so that it is processable by NLP tools.
Outcome: The proposed corpus is larger than existing corpora in terms of size, words and vocabularies.
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology (2026.findings-acl)

Copied to clipboard

Challenge: Dialectal Arabic datasets embody a range of domain, dialect, and quality.
Approach: They propose a framework for automatic speech recognition in dialectal Arabic to address the limited data availability encountered in dialects.
Outcome: The proposed framework provides access to 31 datasets covering 14 dialects to better address the limited data availability encountered in dialectal Arabic speech processing.
An Automatic Learning of an Algerian Dialect Lexicon by using Multilingual Word Embeddings (L18-1)

Copied to clipboard

Challenge: a study on the Algerian Arabic dialect aims to build a lexicon of words written in Arabic or Latin script . multilinguality of the corpus is due to the fact that people use several languages to post comments . stretched letters, misspelled words, emoticons, condensed writing are among the problems .
Approach: They propose to build automatically from a social network an Algerian dialect lexicon.
Outcome: The proposed method leads to a score of 73% on a test lexicon . the study is based on analyzing a lexical corpus of an Algerian dialect .
Dialect-to-Standard Normalization: A Large-Scale Multilingual Evaluation (2023.findings-emnlp)

Copied to clipboard

Challenge: Text normalization is a range of tasks that consist in replacing non-standard spellings with their standard equivalents.
Approach: They introduce dialect-to-standard normalization as a sentence-level character transduction task and provide a large-scale analysis of these methods.
Outcome: The proposed model performs best for Finnish, Swiss German and Slovene while the pre-trained model using full sentences performs the best for Norwegian.
Instruction-Guided Poetry Generation in Arabic and Its Dialects (2026.findings-acl)

Copied to clipboard

Challenge: Existing literature on Arabic poetry has focused on analysis tasks such as interpretation or metadata prediction, e.g., rhyme schemes and titles.
Approach: They propose to use a large-scale instruction-based dataset to generate Arabic poetry based on predefined criteria such as style and rhyme .
Outcome: The proposed model can generate poetry that is aligned with user requirements, based on automated metrics and human evaluation with native Arabic speakers.
On Using Arabic Language Dialects in Recommendation Systems (2025.findings-naacl)

Copied to clipboard

Challenge: Using natural language processing (NLP) to analyze user reviews in recommendation systems is unexplored.
Approach: They propose to integrate Arabic dialects as a signal in recommendation systems by using explicit and implicit approaches.
Outcome: The proposed approach improves recommendation performance and encourages further research in the Arab multicultural world.
Automatic Identification of Maghreb Dialects Using a Dictionary-Based Approach (L18-1)

Copied to clipboard

Challenge: Automatic identification of Arabic dialects in texts is difficult, especially for Maghreb languages and when they are written in Arabic or Latin characters (Arabizi).
Approach: They propose a dictionary-based approach to detect Arabic dialects in texts . they focus on transliteration of Arabicizi into Latin script and code-switching .
Outcome: The proposed approach shows that it is possible to detect dialects in Arabic and Latin scripts.
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs (2025.coling-main)

Copied to clipboard

Challenge: a recent study has found that Arabic is underrepresented in Large Language Models, especially in dialectal variations.
Approach: They propose a benchmark for Arabic Dialect and Cultural Evaluation that evaluates Arabic dialect comprehension and generation.
Outcome: The proposed model outperforms multilingual models on dialect comprehension and generation, but significant challenges persist in dialect identification, generation, and translation.
A description and demonstration of SAFAR framework (2021.eacl-demos)

Copied to clipboard

Challenge: Existing NLP infrastructures are naming them "toolkit", "platform" and "framework" authors present a monolingual framework dedicated to Arabic language .
Approach: They propose a monolingual framework dedicated to Arabic language . they propose namings for existing infrastructures: "toolkit", "platform" and "framework"
Outcome: The proposed framework is dedicated to Arabic language, especially the modern standard Arabic and Moroccan dialect.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations