Anis Charfi, Mabrouka Ben-Sghaier, Andria Samy Raouf Atalla, Raghda Akasheh, Sara Al-Emadi, Wajdi Zaghouani
| Challenge: | Approximately half of the sentences are in Modern Standard Arabic (MSA) for each region, and the other half is in the region’s respective dialect. |
| Approach: | They propose a cross-domain and multi-dialectal stance corpus for Arabic that includes four regions in the Arab World and covers the main Arabic dialect groups. |
| Outcome: | The proposed corpus outperforms the state-of-the-art dataset in stance detection and dialect and dialect classes. |
Similar Papers
The Bahrain Corpus: A Multi-genre Corpus of Bahraini Arabic (2022.lrec-1)
Copied to clipboard
| Challenge: | Various corpora of various sizes and representing different genres, have been created for various Arabic dialects. |
| Approach: | They propose to create a specialized corpus of Bahraini Arabic dialect, which includes written texts as well as transcripts of audio files. |
| Outcome: | The proposed corpus includes 620K words representing the Bahraini Arabic dialect . the annotated corpus is available to support researchers interested in Arabic NLP . |
A Multi-Platform Arabic News Comment Dataset for Offensive Language Detection (2020.lrec-1)
Copied to clipboard
Shammur Absar Chowdhury, Hamdy Mubarak, Ahmed Abdelali, Soon-gyo Jung, Bernard J. Jansen, Joni Salminen
| Challenge: | Social media platforms allow users to engage in conversation with limited accountability, causing hate crimes and mental harm to targeted individuals. |
| Approach: | They propose to make public a new dialectal Arabic news comment dataset . they analyze distinctive lexical content along with the use of emojis in offensive comments . |
| Outcome: | The proposed dataset analyzes offensive language and distinctive lexical content along with the use of emojis on Twitter, Facebook, and YouTube. |
Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs (2026.acl-long)
Copied to clipboard
Abdellah EL Mekki, Samar M. Magdy, Houdaifa Atou, Ruwa AbuHweidi, Baraah Qawasmeh, Omer Nacar, Thikra Al-hibiri, Razan Saadie, Hamzah A. Alsayadi, Nadia Ghezaiel Hammouda, Alshima Mohammed Alkhazimi, Aya Hamod, Al-Yas Yaqoob Al-Ghafri, Wesam El-Sayed, Asila Ismail al Sharji, Mohamad Ballout, Anas Belfathi, Karim Ghaddar, Serry Sibaee, Alaa Aoun, Aeej Mohammed Aseri, Lina Abureesh, Ahlam Bashiti, Majdal Yousef, Abdulaziz Hafiz, Yehdih Mohamed, Emira Hamedtou, Brakehe Emehah, Rahaf Alhamouri, Youssef Nafea, Aya El Aatar, Walid Al-Dhabyani, Emhemed S. Hamed, Sara Shatnawi, Fakhraddin Alwajih, Khalid Elkhidir, Ashwag Alasmari, Abdurrahman Gerrio, Omar Said Alshahri, AbdelRahim A. Elmadany, Ismail Berrada, Amir Azad Adli Al-kathiri, Fadi Zaraket, Mustafa Jarrar, Yahya Mohamed EL Hadj, Hassan Alhuzali, Muhammad Abdul-Mageed
| Challenge: | Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than modern standard Arabic (MSA). |
| Approach: | They propose a large-scale, community-driven, human-translated dataset to bridge this gap . Alexandria covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . |
| Outcome: | The Alexandria dataset covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . Alexandria is a training resource and a rigorous benchmark for evaluating MT and LLMs based on the Alexandria dataset . |
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition (2024.emnlp-main)
Copied to clipboard
Bashar Talafha, Karima Kadaoui, Samar Magdy, Mariem Habiboullah, Chafei Chafei, Ahmed El-Shangiti, Hiba Zayed, Mohamedou Tourad, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, Hour Mohamed, Fakhraddin Alwajih, Abdelrahman Mohamed, Abdellah El Mekki, El Moatez Billah Nagoudi, Benelhadj Saadia, Hamzah Alsayadi, Walid Al-Dhabyani, Sara Shatnawi, Yasir Ech-chammakhy, Amal Makouar, Yousra Berrachedi, Mustafa Jarrar, Shady Shehata, Ismail Berrada, Muhammad Abdul-Mageed
| Challenge: | despite recent advances in speech processing, the majority of world languages and dialects remain uncovered. |
| Approach: | They propose to collect and transcribe a new Arabic dataset for eight dialects . they also develop strong baselines exploiting the new dataset . |
| Outcome: | The proposed dataset covers eight Arabic dialects, including Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni. |
Shami: A Corpus of Levantine Arabic Dialects (L18-1)
Copied to clipboard
| Challenge: | Modern Standard Arabic is the official written language used in education and media . however, the spoken language varies widely across the Arab world . |
| Approach: | They construct a levantine dialect corpus covering data from four dialects spoken in four countries . they describe rules for pre-processing without affecting the meaning so that it is processable by NLP tools. |
| Outcome: | The proposed corpus is larger than existing corpora in terms of size, words and vocabularies. |
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology (2026.findings-acl)
Copied to clipboard
| Challenge: | Dialectal Arabic datasets embody a range of domain, dialect, and quality. |
| Approach: | They propose a framework for automatic speech recognition in dialectal Arabic to address the limited data availability encountered in dialects. |
| Outcome: | The proposed framework provides access to 31 datasets covering 14 dialects to better address the limited data availability encountered in dialectal Arabic speech processing. |
Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks (2025.findings-naacl)
Copied to clipboard
Gagan Bhatia, El Moatez Billah Nagoudi, Abdellah El Mekki, Fakhraddin Alwajih, Muhammad Abdul-Mageed
| Challenge: | In this paper, we introduce a family of embedding models addressing both small-scale and large-scale use cases. |
| Approach: | They propose to use ArabicMTEB to evaluate Arabic text embedding models . they propose to build a benchmark suite that assesses cross-lingual, multi-dialectal, multidomain, and multi-cultural Arabic text embedded models. |
| Outcome: | The proposed models outperform Multilingual-E5-large and Swan-Large in most Arabic tasks while remaining dialectally and culturally aware. |
Konooz: Multi-domain Multi-dialect Corpus for Named Entity Recognition (2025.findings-acl)
Copied to clipboard
| Challenge: | Using the Wojood framework, we compare existing Arabic Named Entity Recognition models with domain and dialect divergence and resource scarcity. |
| Approach: | They propose a multi-dimensional Arabic named entity corpus covering 16 dialects across 10 domains and an annotation scheme using the Wojood guidelines. |
| Outcome: | The proposed model performs better on 16 dialects across 10 domains and 16 domains, while other models struggle with different dialects and domains. |
A description and demonstration of SAFAR framework (2021.eacl-demos)
Copied to clipboard
Karim Bouzoubaa, Younes Jaafar, Driss Namly, Ridouane Tachicart, Rachida Tajmout, Hakima Khamar, Hamid Jaafar, Lhoussain Aouragh, Abdellah Yousfi
| Challenge: | Existing NLP infrastructures are naming them "toolkit", "platform" and "framework" authors present a monolingual framework dedicated to Arabic language . |
| Approach: | They propose a monolingual framework dedicated to Arabic language . they propose namings for existing infrastructures: "toolkit", "platform" and "framework" |
| Outcome: | The proposed framework is dedicated to Arabic language, especially the modern standard Arabic and Moroccan dialect. |
Multilingual Stance Detection in Tweets: The Catalonia Independence Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | stance detection is a method to determine the attitude of a text with respect to a specific topic or claim. |
| Approach: | They propose a multilingual dataset for stance detection in Twitter for the Catalan and Spanish languages. |
| Outcome: | The proposed dataset shows that it is well balanced for multilingual and cross-lingual research. |