| Challenge: | Existing studies of Arabic dialects have focused on blogs and comments on online news sites, but data on other dialects are costly and limited. |
| Approach: | They present a dataset of > 1/4 billion tweets representing a wide range of Arabic dialects. |
| Outcome: | The dataset represents 29 major Arab cities from 10 Arab countries with varying dialects. |
Similar Papers
DART: A Large Dataset of Dialectal Arabic Tweets (L18-1)
Copied to clipboard
| Challenge: | The Arabic language is the fifth most widely spoken language in the world; more than 380 million people speak and write in Arabic. |
| Approach: | They propose to build a large manually-annotated multi-dialect dataset of Arabic tweets that is publicly available. |
| Outcome: | The proposed dataset is well-balanced over five main Arabic dialects: Egyptian, Maghrebi, Levantine, Gulf, and Iraqi. |
Revisiting Common Assumptions about Arabic Dialects in NLP (2025.acl-long)
Copied to clipboard
| Challenge: | Existing assumptions about Arabic dialect variation are not quantitatively verified. |
| Approach: | They extend and analyze Arabic dialects to assess their validity using a multi-label dataset . they find that the assumptions oversimplify reality and are not always accurate . |
| Outcome: | The proposed methods oversimplify reality and are not always accurate, the authors argue . they show that the proposed assumptions oversimply represent reality and may hinder future work . |
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology (2026.findings-acl)
Copied to clipboard
| Challenge: | Dialectal Arabic datasets embody a range of domain, dialect, and quality. |
| Approach: | They propose a framework for automatic speech recognition in dialectal Arabic to address the limited data availability encountered in dialects. |
| Outcome: | The proposed framework provides access to 31 datasets covering 14 dialects to better address the limited data availability encountered in dialectal Arabic speech processing. |
Multi-Dialect Arabic POS Tagging: A CRF Approach (L18-1)
Copied to clipboard
Kareem Darwish, Hamdy Mubarak, Ahmed Abdelali, Mohamed Eldesouki, Younes Samih, Randah Alharbi, Mohammed Attia, Walid Magdy, Laura Kallmeyer
| Challenge: | Existing work on dialectal POS tagging is rather scant with POS tags for most dialects being nonexistent or of limited availability. |
| Approach: | They propose a dataset of POS-tagged Arabic tweets in four major dialects and a tagging guideline for each dialect. |
| Outcome: | The proposed model can tag four different dialects with an average accuracy of 89.3%. |
Fine-Grained Arabic Dialect Identification (C18-1)
Copied to clipboard
| Challenge: | Existing work on Arabic Dialect Identification typically targeted coarse-grained five dialect classes plus Standard Arabic at most (6-way classification). |
| Approach: | They propose to tackle a fine-grained Arabic dialect classification task covering 25 cities from across the Arab World, in addition to Standard Arabic. |
| Outcome: | The proposed task can identify the exact city of a speaker at an accuracy of 67.9% for sentences with an average length of 7 words and reach more than 90% when we consider 16 words. |
ALDi: Quantifying the Arabic Level of Dialectness of Text (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on Dialect Identification (DI) on the sentence level has focused on binary tasks, whereas ALDi treats the task as binary. |
| Approach: | They propose a dataset which contains 127,835 sentences manually labeled with their level of dialectness. |
| Outcome: | The proposed model can identify dialectness on a range of other corpora, providing a more nuanced picture than traditional DI systems. |
On Using Arabic Language Dialects in Recommendation Systems (2025.findings-naacl)
Copied to clipboard
| Challenge: | Using natural language processing (NLP) to analyze user reviews in recommendation systems is unexplored. |
| Approach: | They propose to integrate Arabic dialects as a signal in recommendation systems by using explicit and implicit approaches. |
| Outcome: | The proposed approach improves recommendation performance and encourages further research in the Arab multicultural world. |
Hierarchical Aggregation of Dialectal Data for Arabic Dialect Identification (2022.lrec-1)
Copied to clipboard
| Challenge: | Previous work on Arabic Dialect identification focused on specific dialect levels and labels . since dialectal differences tend to be more subtle relative terms to language differences, the DID task is harder than language identification. |
| Approach: | They propose to define a standard hierarchical schema for Arabic Dialect identification . they map 29 different data sets to this schema and use it to aggregate the data . |
| Outcome: | The proposed schemas and methods are extensible to other languages and dialect groups. |
Arap-Tweet: A Large Multi-Dialect Twitter Corpus for Gender, Age and Language Variety Identification (L18-1)
Copied to clipboard
| Challenge: | Existing corpus of Arabic textual data is limited to English or other European languages. |
| Approach: | They present a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the arab world representing the major Arabic dialectal varieties. |
| Outcome: | The provided corpus will enrich the limited set of available language resources for Arabic and be invaluable enabler for developing author profiling tools and NLP tools for Arabic. |
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition (2024.emnlp-main)
Copied to clipboard
Bashar Talafha, Karima Kadaoui, Samar Magdy, Mariem Habiboullah, Chafei Chafei, Ahmed El-Shangiti, Hiba Zayed, Mohamedou Tourad, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, Hour Mohamed, Fakhraddin Alwajih, Abdelrahman Mohamed, Abdellah El Mekki, El Moatez Billah Nagoudi, Benelhadj Saadia, Hamzah Alsayadi, Walid Al-Dhabyani, Sara Shatnawi, Yasir Ech-chammakhy, Amal Makouar, Yousra Berrachedi, Mustafa Jarrar, Shady Shehata, Ismail Berrada, Muhammad Abdul-Mageed
| Challenge: | despite recent advances in speech processing, the majority of world languages and dialects remain uncovered. |
| Approach: | They propose to collect and transcribe a new Arabic dataset for eight dialects . they also develop strong baselines exploiting the new dataset . |
| Outcome: | The proposed dataset covers eight Arabic dialects, including Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni. |