Challenge: Existing corpus of Arabic textual data is limited to English or other European languages.
Approach: They present a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the arab world representing the major Arabic dialectal varieties.
Outcome: The provided corpus will enrich the limited set of available language resources for Arabic and be invaluable enabler for developing author profiling tools and NLP tools for Arabic.

Similar Papers

DART: A Large Dataset of Dialectal Arabic Tweets (L18-1)

Copied to clipboard

Challenge: The Arabic language is the fifth most widely spoken language in the world; more than 380 million people speak and write in Arabic.
Approach: They propose to build a large manually-annotated multi-dialect dataset of Arabic tweets that is publicly available.
Outcome: The proposed dataset is well-balanced over five main Arabic dialects: Egyptian, Maghrebi, Levantine, Gulf, and Iraqi.
You Tweet What You Speak: A City-Level Dataset of Arabic Dialects (L18-1)

Copied to clipboard

Challenge: Existing studies of Arabic dialects have focused on blogs and comments on online news sites, but data on other dialects are costly and limited.
Approach: They present a dataset of > 1/4 billion tweets representing a wide range of Arabic dialects.
Outcome: The dataset represents 29 major Arab cities from 10 Arab countries with varying dialects.
Multi-Dialect Arabic POS Tagging: A CRF Approach (L18-1)

Copied to clipboard

Challenge: Existing work on dialectal POS tagging is rather scant with POS tags for most dialects being nonexistent or of limited availability.
Approach: They propose a dataset of POS-tagged Arabic tweets in four major dialects and a tagging guideline for each dialect.
Outcome: The proposed model can tag four different dialects with an average accuracy of 89.3%.
DAICT: A Dialectal Arabic Irony Corpus Extracted from Twitter (2020.lrec-1)

Copied to clipboard

Challenge: Current scholarship is yet to reach an agreement on a universal definition of the concept of irony.
Approach: They propose to query Twitter using irony-related hashtags to collect ironic messages which are then manually annotated by two linguists according to their working definition of irony.
Outcome: The proposed corpus will be a valuable resource for developing open domain systems for automatic irony recognition in Arabic and its dialects in social media text.
An Algerian Corpus and an Annotation Platform for Opinion and Emotion Analysis (2020.lrec-1)

Copied to clipboard

Challenge: Currently, there are more than 4 billion Internet users worldwide . the number of social media users in Algeria has tripled over a year .
Approach: They propose a platform for crowdsourcing annotation of tweets at different levels of granularity.
Outcome: The proposed platform can be used to create the largest Algerian dialect subjectivity lexicon of about 9,000 entries.
The Bahrain Corpus: A Multi-genre Corpus of Bahraini Arabic (2022.lrec-1)

Copied to clipboard

Challenge: Various corpora of various sizes and representing different genres, have been created for various Arabic dialects.
Approach: They propose to create a specialized corpus of Bahraini Arabic dialect, which includes written texts as well as transcripts of audio files.
Outcome: The proposed corpus includes 620K words representing the Bahraini Arabic dialect . the annotated corpus is available to support researchers interested in Arabic NLP .
Revisiting Common Assumptions about Arabic Dialects in NLP (2025.acl-long)

Copied to clipboard

Challenge: Existing assumptions about Arabic dialect variation are not quantitatively verified.
Approach: They extend and analyze Arabic dialects to assess their validity using a multi-label dataset . they find that the assumptions oversimplify reality and are not always accurate .
Outcome: The proposed methods oversimplify reality and are not always accurate, the authors argue . they show that the proposed assumptions oversimply represent reality and may hinder future work .
Application and Analysis of a Multi-layered Scheme for Irony on the Italian Twitter Corpus TWITTIRÒ (L18-1)

Copied to clipboard

Challenge: Using a multi-layered scheme for the fine-grained annotation of irony on Italian Twitter is a challenging task to be performed by both human annotators and automatic NLP systems.
Approach: They propose to apply a multi-layered scheme for the fine-grained annotation of irony to an Italian Twitter corpus.
Outcome: The proposed scheme can be validated on Italian irony-laden social media contents and is available in the cross- and multi-lingual perspective.
On Using Arabic Language Dialects in Recommendation Systems (2025.findings-naacl)

Copied to clipboard

Challenge: Using natural language processing (NLP) to analyze user reviews in recommendation systems is unexplored.
Approach: They propose to integrate Arabic dialects as a signal in recommendation systems by using explicit and implicit approaches.
Outcome: The proposed approach improves recommendation performance and encourages further research in the Arab multicultural world.
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs (2025.coling-main)

Copied to clipboard

Challenge: a recent study has found that Arabic is underrepresented in Large Language Models, especially in dialectal variations.
Approach: They propose a benchmark for Arabic Dialect and Cultural Evaluation that evaluates Arabic dialect comprehension and generation.
Outcome: The proposed model outperforms multilingual models on dialect comprehension and generation, but significant challenges persist in dialect identification, generation, and translation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations