Arap-Tweet: A Large Multi-Dialect Twitter Corpus for Gender, Age and Language Variety Identification (L18-1)
Copied to clipboard
| Challenge: | Existing corpus of Arabic textual data is limited to English or other European languages. |
| Approach: | They present a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the arab world representing the major Arabic dialectal varieties. |
| Outcome: | The provided corpus will enrich the limited set of available language resources for Arabic and be invaluable enabler for developing author profiling tools and NLP tools for Arabic. |
Similar Papers
DART: A Large Dataset of Dialectal Arabic Tweets (L18-1)
Copied to clipboard
| Challenge: | The Arabic language is the fifth most widely spoken language in the world; more than 380 million people speak and write in Arabic. |
| Approach: | They propose to build a large manually-annotated multi-dialect dataset of Arabic tweets that is publicly available. |
| Outcome: | The proposed dataset is well-balanced over five main Arabic dialects: Egyptian, Maghrebi, Levantine, Gulf, and Iraqi. |
You Tweet What You Speak: A City-Level Dataset of Arabic Dialects (L18-1)
Copied to clipboard
| Challenge: | Existing studies of Arabic dialects have focused on blogs and comments on online news sites, but data on other dialects are costly and limited. |
| Approach: | They present a dataset of > 1/4 billion tweets representing a wide range of Arabic dialects. |
| Outcome: | The dataset represents 29 major Arab cities from 10 Arab countries with varying dialects. |
Multi-Dialect Arabic POS Tagging: A CRF Approach (L18-1)
Copied to clipboard
Kareem Darwish, Hamdy Mubarak, Ahmed Abdelali, Mohamed Eldesouki, Younes Samih, Randah Alharbi, Mohammed Attia, Walid Magdy, Laura Kallmeyer
| Challenge: | Existing work on dialectal POS tagging is rather scant with POS tags for most dialects being nonexistent or of limited availability. |
| Approach: | They propose a dataset of POS-tagged Arabic tweets in four major dialects and a tagging guideline for each dialect. |
| Outcome: | The proposed model can tag four different dialects with an average accuracy of 89.3%. |
DAICT: A Dialectal Arabic Irony Corpus Extracted from Twitter (2020.lrec-1)
Copied to clipboard
| Challenge: | Current scholarship is yet to reach an agreement on a universal definition of the concept of irony. |
| Approach: | They propose to query Twitter using irony-related hashtags to collect ironic messages which are then manually annotated by two linguists according to their working definition of irony. |
| Outcome: | The proposed corpus will be a valuable resource for developing open domain systems for automatic irony recognition in Arabic and its dialects in social media text. |
An Algerian Corpus and an Annotation Platform for Opinion and Emotion Analysis (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, there are more than 4 billion Internet users worldwide . the number of social media users in Algeria has tripled over a year . |
| Approach: | They propose a platform for crowdsourcing annotation of tweets at different levels of granularity. |
| Outcome: | The proposed platform can be used to create the largest Algerian dialect subjectivity lexicon of about 9,000 entries. |
The Bahrain Corpus: A Multi-genre Corpus of Bahraini Arabic (2022.lrec-1)
Copied to clipboard
| Challenge: | Various corpora of various sizes and representing different genres, have been created for various Arabic dialects. |
| Approach: | They propose to create a specialized corpus of Bahraini Arabic dialect, which includes written texts as well as transcripts of audio files. |
| Outcome: | The proposed corpus includes 620K words representing the Bahraini Arabic dialect . the annotated corpus is available to support researchers interested in Arabic NLP . |
Revisiting Common Assumptions about Arabic Dialects in NLP (2025.acl-long)
Copied to clipboard
| Challenge: | Existing assumptions about Arabic dialect variation are not quantitatively verified. |
| Approach: | They extend and analyze Arabic dialects to assess their validity using a multi-label dataset . they find that the assumptions oversimplify reality and are not always accurate . |
| Outcome: | The proposed methods oversimplify reality and are not always accurate, the authors argue . they show that the proposed assumptions oversimply represent reality and may hinder future work . |
Application and Analysis of a Multi-layered Scheme for Irony on the Italian Twitter Corpus TWITTIRÒ (L18-1)
Copied to clipboard
| Challenge: | Using a multi-layered scheme for the fine-grained annotation of irony on Italian Twitter is a challenging task to be performed by both human annotators and automatic NLP systems. |
| Approach: | They propose to apply a multi-layered scheme for the fine-grained annotation of irony to an Italian Twitter corpus. |
| Outcome: | The proposed scheme can be validated on Italian irony-laden social media contents and is available in the cross- and multi-lingual perspective. |
On Using Arabic Language Dialects in Recommendation Systems (2025.findings-naacl)
Copied to clipboard
| Challenge: | Using natural language processing (NLP) to analyze user reviews in recommendation systems is unexplored. |
| Approach: | They propose to integrate Arabic dialects as a signal in recommendation systems by using explicit and implicit approaches. |
| Outcome: | The proposed approach improves recommendation performance and encourages further research in the Arab multicultural world. |
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs (2025.coling-main)
Copied to clipboard
Basel Mousi, Nadir Durrani, Fatema Ahmad, Md. Arid Hasan, Maram Hasanain, Tameem Kabbani, Fahim Dalvi, Shammur Absar Chowdhury, Firoj Alam
| Challenge: | a recent study has found that Arabic is underrepresented in Large Language Models, especially in dialectal variations. |
| Approach: | They propose a benchmark for Arabic Dialect and Cultural Evaluation that evaluates Arabic dialect comprehension and generation. |
| Outcome: | The proposed model outperforms multilingual models on dialect comprehension and generation, but significant challenges persist in dialect identification, generation, and translation. |