Papers by Wajdi Zaghouani
DAICT: A Dialectal Arabic Irony Corpus Extracted from Twitter (2020.lrec-1)
Copied to clipboard
| Challenge: | Current scholarship is yet to reach an agreement on a universal definition of the concept of irony. |
| Approach: | They propose to query Twitter using irony-related hashtags to collect ironic messages which are then manually annotated by two linguists according to their working definition of irony. |
| Outcome: | The proposed corpus will be a valuable resource for developing open domain systems for automatic irony recognition in Arabic and its dialects in social media text. |
The MADAR Arabic Dialect Corpus and Lexicon (L18-1)
Copied to clipboard
Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, Kemal Oflazer
| Challenge: | Using a corpus of 25 Arabic city dialects and a lexicon of 1,045 concepts, we study 25 cities in a travel domain . focus on cities opens new avenues for research from dialectology to dialect identification and machine translation. |
| Approach: | They present two Arabic language resources that are part of the Multi Arabic Dialect Applications and Resources project. |
| Outcome: | The proposed resources are the first of their kind in terms of their coverage and fine granularity. |
QCAW 1.0: Building a Qatari Corpus of Student Argumentative Writing (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies have highlighted the importance of and need to create learner corpora. |
| Approach: | They propose to create a Qatari corpus of argumentative writing (QCAW) the corpus contains 200,000 tokens of argumentation written by Qatari university students . |
| Outcome: | The QCAW contains 195 essays written by 195 students, 159 females and 36 males. |
Unified Guidelines and Resources for Arabic Dialect Orthography (L18-1)
Copied to clipboard
Nizar Habash, Fadhl Eryani, Salam Khalifa, Owen Rambow, Dana Abdulrahim, Alexander Erdmann, Reem Faraj, Wajdi Zaghouani, Houda Bouamor, Nasser Zalmout, Sara Hassan, Faisal Al-Shargi, Sakhar Alkhereyf, Basma Abdulkareem, Ramy Eskander, Mohammad Salameh, Hind Saddiki
| Challenge: | Existing efforts to conventionalize the dialectal orthography of Arabic have focused on specific dialects and made ad hoc decisions. |
| Approach: | They propose a set of guidelines and meta-guidelines for conventional orthography of Arabic dialects . they apply them to 28 Arab city dialects from Rabat to Muscat . |
| Outcome: | The proposed guidelines and resources are being used by three large Arabic dialect processing projects in three universities. |
MADARi: A Web Interface for Joint Arabic Morphological Annotation and Spelling Correction (L18-1)
Copied to clipboard
| Challenge: | Standard Arabic morphology is rich, but Arabic dialects introduce more complexity. |
| Approach: | They propose a joint morphological annotation and spelling correction system for Arabic texts . they propose morphology tools that can be used to help with productivity . |
| Outcome: | The proposed system is based on a standard and dialectal Arabic text. |
So Hateful! Building a Multi-Label Hate Speech Annotated Arabic Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | Social media enables widespread propagation of hate speech targeting groups based on ethnicity, religion, or other characteristics. |
| Approach: | They analyze 70,000 Arabic tweets to identify hate speech patterns and train models . 15% of tweets contain offensive language while 6% have hate speech . authors hope to prevent spread of hateful content on social media platforms . |
| Outcome: | The analysis of 70,000 Arabic tweets shows that 15% of tweets contain offensive language while 6% have hate speech . 10% of tweet provide verifiable factual claims, and 7% are deemed important . |
A Multi-Task Learning Framework for Modeling Engagement and Topic-Sensitive Responses in Arabic Women’s Discourse (2026.findings-eacl)
Copied to clipboard
| Challenge: | a corpus of 158k arab Facebook posts spanning women's rights, gender debates, and economic empowerment reveals patterns of public opinion that vary dramatically across regional and cultural contexts. |
| Approach: | They propose a multi-task learning framework that learns audience reaction classification and engagement magnitude regression and non-engagement detection. |
| Outcome: | The proposed model achieves a test macro-F1 of 72.4 and weighted-F1. It measures 158k posts across gender issues, legal rights advocacy, gender identity discussions, and economic empowerment. |
Arap-Tweet: A Large Multi-Dialect Twitter Corpus for Gender, Age and Language Variety Identification (L18-1)
Copied to clipboard
| Challenge: | Existing corpus of Arabic textual data is limited to English or other European languages. |
| Approach: | They present a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the arab world representing the major Arabic dialectal varieties. |
| Outcome: | The provided corpus will enrich the limited set of available language resources for Arabic and be invaluable enabler for developing author profiling tools and NLP tools for Arabic. |
Fighting the COVID-19 Infodemic: Modeling the Perspective of Journalists, Fact-Checkers, Social Media Platforms, Policy Makers, and the Society (2021.findings-emnlp)
Copied to clipboard
Firoj Alam, Shaden Shaar, Fahim Dalvi, Hassan Sajjad, Alex Nikolov, Hamdy Mubarak, Giovanni Da San Martino, Ahmed Abdelali, Nadir Durrani, Kareem Darwish, Abdulaziz Al-Homaid, Wajdi Zaghouani, Tommaso Caselli, Gijs Danoe, Friso Stolk, Britt Bruntink, Preslav Nakov
| Challenge: | a dataset of 16K manually annotated tweets is used to analyze disinformation . the democratic nature of social media has raised questions about the quality and the factuality of the information that is shared on these platforms. |
| Approach: | They use a dataset of manually annotated tweets to analyze COVID-19 disinformation . they show that tweets contain fake cures, rumors, conspiracy theories and xenophobia . |
| Outcome: | The proposed dataset shows that it is useful in monolingual vs. multilingual settings. |
MARASTA: A Multi-dialectal Arabic Cross-domain Stance Corpus (2024.lrec-main)
Copied to clipboard
Anis Charfi, Mabrouka Ben-Sghaier, Andria Samy Raouf Atalla, Raghda Akasheh, Sara Al-Emadi, Wajdi Zaghouani
| Challenge: | Approximately half of the sentences are in Modern Standard Arabic (MSA) for each region, and the other half is in the region’s respective dialect. |
| Approach: | They propose a cross-domain and multi-dialectal stance corpus for Arabic that includes four regions in the Arab World and covers the main Arabic dialect groups. |
| Outcome: | The proposed corpus outperforms the state-of-the-art dataset in stance detection and dialect and dialect classes. |