Towards Responsible Natural Language Annotation for the Varieties of Arabic (2022.findings-acl)
Copied to clipboard
| Challenge: | In NLP, there is a tendency to aim for broader coverage, often overlooking cultural and (socio)linguistic nuance. |
| Approach: | They propose a playbook for responsible dataset creation for polyglossic, multidialectal languages . they focus on Arabic annotation of social media content as an example . |
| Outcome: | The proposed model is based on Arabic annotation of social media content. |
Similar Papers
Geo-Cultural Representation and Inclusion in Language Technologies (2024.lrec-tutorials)
Copied to clipboard
| Challenge: | audi et al.: training and evaluation of language models rely on semi-structured data that is annotated by humans . e-learning tools do not integrate rich and diverse community perspectives into language technologies . |
| Approach: | They will examine how different socio-cultural perspectives influence what is taken as ground truth by models. |
| Outcome: | This tutorial examines how different socio-cultural perspectives influence representations of global concepts. |
Masader: Metadata Sourcing for Arabic Text and Speech Data Resources (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, there is no online catalogue for Arabic datasets with annotated attributes . this paper aims to identify the publicly available Arabic dataset and provide a catalogue of them to researchers. |
| Approach: | They propose to create the largest public catalogue for Arabic NLP datasets with 25 attributes and a metadata annotation strategy that could be extended to other languages. |
| Outcome: | The proposed approach could be extended to other languages and regions. |
Challenges and Strategies in Cross-Cultural NLP (2022.acl-long)
Copied to clipboard
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard
| Challenge: | Various efforts have been made to accommodate linguistic diversity and serve speakers of many different languages. |
| Approach: | They propose a framework to examine cultural differences in NLP to better serve users . they argue that cultural knowledge, preferences and values can affect NLP practices . |
| Outcome: | The proposed framework examines how cultural knowledge, preferences and values can affect NLP practices. |
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)
Copied to clipboard
| Challenge: | Language is a powerful means of communication and should be regarded as more than just a collection of tokens. |
| Approach: | They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices. |
| Outcome: | The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices. |
Modeling the Sacred: Considerations when Using Religious Texts in Natural Language Processing (2024.findings-naacl)
Copied to clipboard
| Challenge: | This paper concerns the use of religious texts in natural language processing (NLP) religious texts are expressions of culturally important values, and machine learning models reproduce cultural values encoded in training data. |
| Approach: | They argue that NLP's use of religious texts raises considerations beyond model biases . authors argue that religious texts are culturally important and are often used by researchers . |
| Outcome: | The proposed method repurposes translations from their original uses and motivations, and raises considerations beyond model biases. |
A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions (2025.coling-main)
Copied to clipboard
| Challenge: | Code-switching (CSW) is a common linguistic phenomenon in multilingual societies . current literature on CSW in the arab world is limited to the Arabic language . |
| Approach: | They present a review of the literature in the field of code-switched Arabic NLP . they propose recommendations for future research . |
| Outcome: | This review provides a broad perspective on the current literature in the field of code-switched Arabic NLP . it also provides recommendations for future research . |
Arabic Natural Language Processing (2022.emnlp-tutorials)
Copied to clipboard
| Challenge: | This tutorial provides background information for system developers and researchers working with Arabic in its various forms. |
| Approach: | This tutorial provides the necessary background information for working with Arabic in its various forms. |
| Outcome: | This tutorial will explain various Arabic linguistic phenomena and review the state-of-the-art in Arabic processing. |
Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs (2025.acl-long)
Copied to clipboard
Fakhraddin Alwajih, Abdellah El Mekki, Samar Mohamed Magdy, AbdelRahim A. Elmadany, Omer Nacar, El Moatez Billah Nagoudi, Reem Abdel-Salam, Hanin Atwany, Youssef Nafea, Abdulfattah Mohammed Yahya, Rahaf Alhamouri, Hamzah A. Alsayadi, Hiba Zayed, Sara Shatnawi, Serry Sibaee, Yasir Ech-chammakhy, Walid Al-Dhabyani, Marwa Mohamed Ali, Imen Jarraya, Ahmed Oumar El-Shangiti, Aisha Alraeesi, Mohammed Anwar AL-Ghrawi, Abdulrahman S. Al-Batati, Elgizouli Mohamed, Noha Taha Elgindi, Muhammed Saeed, Houdaifa Atou, Issam Ait Yahia, Abdelhak Bouayad, Mohammed Machrouh, Amal Makouar, Dania Alkawi, Mukhtar Mohamed, Safaa Taher Abdelfadil, Amine Ziad Ounnoughene, Anfel Rouabhia, Rwaa Assi, Ahmed Sorkatti, Mohamedou Cheikh Tourad, Anis Koubaa, Ismail Berrada, Mustafa Jarrar, Shady Shehata, Muhammad Abdul-Mageed
| Challenge: | a year-long community-driven project covering all 22 Arab countries evaluates the cultural and dialectal capabilities of large language models. |
| Approach: | They propose a project to evaluate the cultural and dialectal capabilities of large language models. |
| Outcome: | The project evaluates the cultural and dialectal capabilities of several frontier LLMs. |
The KIND Dataset: A Social Collaboration Approach for Nuanced Dialect Data Collection (2024.eacl-srw)
Copied to clipboard
Asma Yamani, Raghad Alziyady, Reem AlYami, Salma Albelali, Leina Albelali, Jawharah Almulhim, Amjad Alsulami, Motaz Alfarraj, Rabeah Al-Zaidy
| Challenge: | Nuanced dialects are a linguistic variant that pose several challenges for NLP models and techniques. |
| Approach: | They propose an approach to collect high quality Arabic dialect data by social collaboration . they use short texts to collect Arabic dialects and a KIND corpus . |
| Outcome: | The proposed approach is based on the KIND corpus of Arabic dialect data . it provides a high quality dataset and is versatile enough to be multipurpose . |
NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities (2025.emnlp-main)
Copied to clipboard
| Challenge: | Current research directions rely on synthetic data generated by translating English corpora, which often fails to represent the cultural heritage and values of local communities. |
| Approach: | They propose a method to create and retrieve pre-training data tailored to a specific community . they use Egyptian and Moroccan dialects as testbeds to test their understanding . |
| Outcome: | The proposed method outperforms existing Arabic-aware LLMs and performs on par with larger models. |