| Challenge: | Existing work on Arabic Dialect Identification typically targeted coarse-grained five dialect classes plus Standard Arabic at most (6-way classification). |
| Approach: | They propose to tackle a fine-grained Arabic dialect classification task covering 25 cities from across the Arab World, in addition to Standard Arabic. |
| Outcome: | The proposed task can identify the exact city of a speaker at an accuracy of 67.9% for sentences with an average length of 7 words and reach more than 90% when we consider 16 words. |
Similar Papers
You Tweet What You Speak: A City-Level Dataset of Arabic Dialects (L18-1)
Copied to clipboard
| Challenge: | Existing studies of Arabic dialects have focused on blogs and comments on online news sites, but data on other dialects are costly and limited. |
| Approach: | They present a dataset of > 1/4 billion tweets representing a wide range of Arabic dialects. |
| Outcome: | The dataset represents 29 major Arab cities from 10 Arab countries with varying dialects. |
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology (2026.findings-acl)
Copied to clipboard
| Challenge: | Dialectal Arabic datasets embody a range of domain, dialect, and quality. |
| Approach: | They propose a framework for automatic speech recognition in dialectal Arabic to address the limited data availability encountered in dialects. |
| Outcome: | The proposed framework provides access to 31 datasets covering 14 dialects to better address the limited data availability encountered in dialectal Arabic speech processing. |
Revisiting Common Assumptions about Arabic Dialects in NLP (2025.acl-long)
Copied to clipboard
| Challenge: | Existing assumptions about Arabic dialect variation are not quantitatively verified. |
| Approach: | They extend and analyze Arabic dialects to assess their validity using a multi-label dataset . they find that the assumptions oversimplify reality and are not always accurate . |
| Outcome: | The proposed methods oversimplify reality and are not always accurate, the authors argue . they show that the proposed assumptions oversimply represent reality and may hinder future work . |
The MADAR Arabic Dialect Corpus and Lexicon (L18-1)
Copied to clipboard
Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, Kemal Oflazer
| Challenge: | Using a corpus of 25 Arabic city dialects and a lexicon of 1,045 concepts, we study 25 cities in a travel domain . focus on cities opens new avenues for research from dialectology to dialect identification and machine translation. |
| Approach: | They present two Arabic language resources that are part of the Multi Arabic Dialect Applications and Resources project. |
| Outcome: | The proposed resources are the first of their kind in terms of their coverage and fine granularity. |
Hierarchical Aggregation of Dialectal Data for Arabic Dialect Identification (2022.lrec-1)
Copied to clipboard
| Challenge: | Previous work on Arabic Dialect identification focused on specific dialect levels and labels . since dialectal differences tend to be more subtle relative terms to language differences, the DID task is harder than language identification. |
| Approach: | They propose to define a standard hierarchical schema for Arabic Dialect identification . they map 29 different data sets to this schema and use it to aggregate the data . |
| Outcome: | The proposed schemas and methods are extensible to other languages and dialect groups. |
ALDi: Quantifying the Arabic Level of Dialectness of Text (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on Dialect Identification (DI) on the sentence level has focused on binary tasks, whereas ALDi treats the task as binary. |
| Approach: | They propose a dataset which contains 127,835 sentences manually labeled with their level of dialectness. |
| Outcome: | The proposed model can identify dialectness on a range of other corpora, providing a more nuanced picture than traditional DI systems. |
Text and Speech-based Tunisian Arabic Sub-Dialects Identification (2020.lrec-1)
Copied to clipboard
| Challenge: | Dialect IDentification is a difficult task when it is about the identification of dialects belonging to the same country. |
| Approach: | They present results on a dialect classification task covering four sub-dialects spoken in Tunisia using a spoken corpus of 1673 utterances. |
| Outcome: | The proposed system achieves an F-1 score of 93.75% while the F-1 is limited to 54.16% using text-based DID on the same test set. |
A Spelling Correction Corpus for Multiple Arabic Dialects (2020.lrec-1)
Copied to clipboard
| Challenge: | Arabic dialects are non-standard varieties of Arabic commonly spoken across the Arab world, but lack standard orthographies. |
| Approach: | They present a corpus of 10,000 sentences from five Arabic city dialects represented in the Conventional Orthography for Dialectal Arabic (CODA) they use a bootstrapping technique to speed up annotation and compare similarity between dialects before and after CODA annotation. |
| Outcome: | The proposed method speeds up the annotation process and shows similarity between the dialects before and after CODA annotation. |
Unified Guidelines and Resources for Arabic Dialect Orthography (L18-1)
Copied to clipboard
Nizar Habash, Fadhl Eryani, Salam Khalifa, Owen Rambow, Dana Abdulrahim, Alexander Erdmann, Reem Faraj, Wajdi Zaghouani, Houda Bouamor, Nasser Zalmout, Sara Hassan, Faisal Al-Shargi, Sakhar Alkhereyf, Basma Abdulkareem, Ramy Eskander, Mohammad Salameh, Hind Saddiki
| Challenge: | Existing efforts to conventionalize the dialectal orthography of Arabic have focused on specific dialects and made ad hoc decisions. |
| Approach: | They propose a set of guidelines and meta-guidelines for conventional orthography of Arabic dialects . they apply them to 28 Arab city dialects from Rabat to Muscat . |
| Outcome: | The proposed guidelines and resources are being used by three large Arabic dialect processing projects in three universities. |
ADIDA: Automatic Dialect Identification for Arabic (N19-4)
Copied to clipboard
| Challenge: | Demo paper describes a web-based system for automatic dialect identification for Arabic text. |
| Approach: | They present a web-based system for automatic dialect identification for Arabic text that distinguishes between 25 Arab cities and Modern Standard Arabic. |
| Outcome: | The proposed system distinguishes among the dialects of 25 Arab cities (from Rabat to Muscat) and Modern Standard Arabic (MSA). |