Challenge: In NLP, there is a tendency to aim for broader coverage, often overlooking cultural and (socio)linguistic nuance.
Approach: They propose a playbook for responsible dataset creation for polyglossic, multidialectal languages . they focus on Arabic annotation of social media content as an example .
Outcome: The proposed model is based on Arabic annotation of social media content.

Similar Papers

Geo-Cultural Representation and Inclusion in Language Technologies (2024.lrec-tutorials)

Copied to clipboard

Challenge: audi et al.: training and evaluation of language models rely on semi-structured data that is annotated by humans . e-learning tools do not integrate rich and diverse community perspectives into language technologies .
Approach: They will examine how different socio-cultural perspectives influence what is taken as ground truth by models.
Outcome: This tutorial examines how different socio-cultural perspectives influence representations of global concepts.
Masader: Metadata Sourcing for Arabic Text and Speech Data Resources (2022.lrec-1)

Copied to clipboard

Challenge: Currently, there is no online catalogue for Arabic datasets with annotated attributes . this paper aims to identify the publicly available Arabic dataset and provide a catalogue of them to researchers.
Approach: They propose to create the largest public catalogue for Arabic NLP datasets with 25 attributes and a metadata annotation strategy that could be extended to other languages.
Outcome: The proposed approach could be extended to other languages and regions.
Challenges and Strategies in Cross-Cultural NLP (2022.acl-long)

Copied to clipboard

Challenge: Various efforts have been made to accommodate linguistic diversity and serve speakers of many different languages.
Approach: They propose a framework to examine cultural differences in NLP to better serve users . they argue that cultural knowledge, preferences and values can affect NLP practices .
Outcome: The proposed framework examines how cultural knowledge, preferences and values can affect NLP practices.
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)

Copied to clipboard

Challenge: Language is a powerful means of communication and should be regarded as more than just a collection of tokens.
Approach: They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices.
Outcome: The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices.
Modeling the Sacred: Considerations when Using Religious Texts in Natural Language Processing (2024.findings-naacl)

Copied to clipboard

Challenge: This paper concerns the use of religious texts in natural language processing (NLP) religious texts are expressions of culturally important values, and machine learning models reproduce cultural values encoded in training data.
Approach: They argue that NLP's use of religious texts raises considerations beyond model biases . authors argue that religious texts are culturally important and are often used by researchers .
Outcome: The proposed method repurposes translations from their original uses and motivations, and raises considerations beyond model biases.
A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions (2025.coling-main)

Copied to clipboard

Challenge: Code-switching (CSW) is a common linguistic phenomenon in multilingual societies . current literature on CSW in the arab world is limited to the Arabic language .
Approach: They present a review of the literature in the field of code-switched Arabic NLP . they propose recommendations for future research .
Outcome: This review provides a broad perspective on the current literature in the field of code-switched Arabic NLP . it also provides recommendations for future research .
Arabic Natural Language Processing (2022.emnlp-tutorials)

Copied to clipboard

Challenge: This tutorial provides background information for system developers and researchers working with Arabic in its various forms.
Approach: This tutorial provides the necessary background information for working with Arabic in its various forms.
Outcome: This tutorial will explain various Arabic linguistic phenomena and review the state-of-the-art in Arabic processing.
The KIND Dataset: A Social Collaboration Approach for Nuanced Dialect Data Collection (2024.eacl-srw)

Copied to clipboard

Challenge: Nuanced dialects are a linguistic variant that pose several challenges for NLP models and techniques.
Approach: They propose an approach to collect high quality Arabic dialect data by social collaboration . they use short texts to collect Arabic dialects and a KIND corpus .
Outcome: The proposed approach is based on the KIND corpus of Arabic dialect data . it provides a high quality dataset and is versatile enough to be multipurpose .
NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities (2025.emnlp-main)

Copied to clipboard

Challenge: Current research directions rely on synthetic data generated by translating English corpora, which often fails to represent the cultural heritage and values of local communities.
Approach: They propose a method to create and retrieve pre-training data tailored to a specific community . they use Egyptian and Moroccan dialects as testbeds to test their understanding .
Outcome: The proposed method outperforms existing Arabic-aware LLMs and performs on par with larger models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations