Fine-Grained Arabic Dialect Identification (C18-1)

Copied to clipboard

Challenge: Existing work on Arabic Dialect Identification typically targeted coarse-grained five dialect classes plus Standard Arabic at most (6-way classification).
Approach: They propose to tackle a fine-grained Arabic dialect classification task covering 25 cities from across the Arab World, in addition to Standard Arabic.
Outcome: The proposed task can identify the exact city of a speaker at an accuracy of 67.9% for sentences with an average length of 7 words and reach more than 90% when we consider 16 words.

Similar Papers

You Tweet What You Speak: A City-Level Dataset of Arabic Dialects (L18-1)

Copied to clipboard

Challenge: Existing studies of Arabic dialects have focused on blogs and comments on online news sites, but data on other dialects are costly and limited.
Approach: They present a dataset of > 1/4 billion tweets representing a wide range of Arabic dialects.
Outcome: The dataset represents 29 major Arab cities from 10 Arab countries with varying dialects.
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology (2026.findings-acl)

Copied to clipboard

Challenge: Dialectal Arabic datasets embody a range of domain, dialect, and quality.
Approach: They propose a framework for automatic speech recognition in dialectal Arabic to address the limited data availability encountered in dialects.
Outcome: The proposed framework provides access to 31 datasets covering 14 dialects to better address the limited data availability encountered in dialectal Arabic speech processing.
Revisiting Common Assumptions about Arabic Dialects in NLP (2025.acl-long)

Copied to clipboard

Challenge: Existing assumptions about Arabic dialect variation are not quantitatively verified.
Approach: They extend and analyze Arabic dialects to assess their validity using a multi-label dataset . they find that the assumptions oversimplify reality and are not always accurate .
Outcome: The proposed methods oversimplify reality and are not always accurate, the authors argue . they show that the proposed assumptions oversimply represent reality and may hinder future work .
The MADAR Arabic Dialect Corpus and Lexicon (L18-1)

Copied to clipboard

Challenge: Using a corpus of 25 Arabic city dialects and a lexicon of 1,045 concepts, we study 25 cities in a travel domain . focus on cities opens new avenues for research from dialectology to dialect identification and machine translation.
Approach: They present two Arabic language resources that are part of the Multi Arabic Dialect Applications and Resources project.
Outcome: The proposed resources are the first of their kind in terms of their coverage and fine granularity.
Hierarchical Aggregation of Dialectal Data for Arabic Dialect Identification (2022.lrec-1)

Copied to clipboard

Challenge: Previous work on Arabic Dialect identification focused on specific dialect levels and labels . since dialectal differences tend to be more subtle relative terms to language differences, the DID task is harder than language identification.
Approach: They propose to define a standard hierarchical schema for Arabic Dialect identification . they map 29 different data sets to this schema and use it to aggregate the data .
Outcome: The proposed schemas and methods are extensible to other languages and dialect groups.
ALDi: Quantifying the Arabic Level of Dialectness of Text (2023.emnlp-main)

Copied to clipboard

Challenge: Existing work on Dialect Identification (DI) on the sentence level has focused on binary tasks, whereas ALDi treats the task as binary.
Approach: They propose a dataset which contains 127,835 sentences manually labeled with their level of dialectness.
Outcome: The proposed model can identify dialectness on a range of other corpora, providing a more nuanced picture than traditional DI systems.
Text and Speech-based Tunisian Arabic Sub-Dialects Identification (2020.lrec-1)

Copied to clipboard

Challenge: Dialect IDentification is a difficult task when it is about the identification of dialects belonging to the same country.
Approach: They present results on a dialect classification task covering four sub-dialects spoken in Tunisia using a spoken corpus of 1673 utterances.
Outcome: The proposed system achieves an F-1 score of 93.75% while the F-1 is limited to 54.16% using text-based DID on the same test set.
A Spelling Correction Corpus for Multiple Arabic Dialects (2020.lrec-1)

Copied to clipboard

Challenge: Arabic dialects are non-standard varieties of Arabic commonly spoken across the Arab world, but lack standard orthographies.
Approach: They present a corpus of 10,000 sentences from five Arabic city dialects represented in the Conventional Orthography for Dialectal Arabic (CODA) they use a bootstrapping technique to speed up annotation and compare similarity between dialects before and after CODA annotation.
Outcome: The proposed method speeds up the annotation process and shows similarity between the dialects before and after CODA annotation.
Unified Guidelines and Resources for Arabic Dialect Orthography (L18-1)

Copied to clipboard

Challenge: Existing efforts to conventionalize the dialectal orthography of Arabic have focused on specific dialects and made ad hoc decisions.
Approach: They propose a set of guidelines and meta-guidelines for conventional orthography of Arabic dialects . they apply them to 28 Arab city dialects from Rabat to Muscat .
Outcome: The proposed guidelines and resources are being used by three large Arabic dialect processing projects in three universities.
ADIDA: Automatic Dialect Identification for Arabic (N19-4)

Copied to clipboard

Challenge: Demo paper describes a web-based system for automatic dialect identification for Arabic text.
Approach: They present a web-based system for automatic dialect identification for Arabic text that distinguishes between 25 Arab cities and Modern Standard Arabic.
Outcome: The proposed system distinguishes among the dialects of 25 Arab cities (from Rabat to Muscat) and Modern Standard Arabic (MSA).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations