Challenge: a dataset of social media texts addressing LGBTQIA+ individuals is presented in this paper . the dataset is based on two sources in italian: Facebook and Twitter .
Approach: They describe a dataset composed of two sub-corpora from two different sources in Italian . the dataset includes social media texts regarding LGBTQIA+ individuals, behaviors, ideology and events .
Outcome: The QUEEREOTYPES dataset includes social media texts regarding LGBTQIA+ individuals, behaviors, ideology and events.

Similar Papers

An Italian Twitter Corpus of Hate Speech against Immigrants (L18-1)

Copied to clipboard

Challenge: a recent study has annotated 6,000 tweets for hate speech against immigrants . the annotation scheme was designed to account for the multiplicity of factors that can contribute to the definition of a hate speech notion .
Approach: They describe a Twitter corpus annotated for hate speech against immigrants . they propose a scheme that includes aggressiveness, offensiveness, irony, stereotype and intensity .
Outcome: The proposed annotation scheme includes aggressiveness, offensiveness, irony, stereotype, intensity and (on an experimental basis) intensity.
Application and Analysis of a Multi-layered Scheme for Irony on the Italian Twitter Corpus TWITTIRÒ (L18-1)

Copied to clipboard

Challenge: Using a multi-layered scheme for the fine-grained annotation of irony on Italian Twitter is a challenging task to be performed by both human annotators and automatic NLP systems.
Approach: They propose to apply a multi-layered scheme for the fine-grained annotation of irony to an Italian Twitter corpus.
Outcome: The proposed scheme can be validated on Italian irony-laden social media contents and is available in the cross- and multi-lingual perspective.
A Multilingual Dataset of Racial Stereotypes in Social Media Conversational Threads (2023.findings-eacl)

Copied to clipboard

Challenge: a new corpus-based study addresses racial stereotypes in social media conversations . a multilingual corpus of rhs is used to investigate how they are spread .
Approach: They propose a corpus-based method for multilingual racial stereotype identification in social media conversational threads.
Outcome: The proposed method sheds light on how racial hoaxes are spread and allows identification of negative stereotypes that reinforce them.
Adaptive Data Collection for Latin-American Community-sourced Evaluation of Stereotypes (LACES) (2026.findings-acl)

Copied to clipboard

Challenge: a geo-cultural gap in NLP evaluation hinders evaluation of societal biases . authors propose a new method to collect stereotypes from large language models .
Approach: They propose a new method that integrates sourcing and validation of existing data into a single workflow.
Outcome: The proposed method improves LACES by integrating new stereotype entries and validation of existing data.
Your Stereotypical Mileage May Vary: Practical Challenges of Evaluating Biases in Multiple Languages and Cultural Contexts (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies have identified a gap in the availability of tools and resources to study bias in languages other than English and social contexts outside the north of America.
Approach: They use stereotypes to build a corpus of sentence pairs that cover biases in seven cultural contexts.
Outcome: The proposed resource covers a wide range of languages and cultural settings . it favors sentences that express stereotypes in most bias categories .
Intersectional Stereotypes in Large Language Models: Dataset and Analysis (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on intersectional stereotypes focus on broader, individual categories . current studies focus on single-group stereotypes, such as racial bias against African Americans .
Approach: They propose to use a dataset of intersectional stereotypes curated with the ChatGPT model to analyze propagation in three contemporary LLMs.
Outcome: The proposed dataset enables analysis of stereotype propagation in three contemporary LLMs.
Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Existing work on document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes.
Approach: They assemble and publish a multilingual Twitter corpus for the task of hate speech detection using inferred author demographic factors.
Outcome: The results show that the classifiers learn human biases and can be discriminatory towards certain demographic groups.
SAFARI: A Community-Engaged Approach and Dataset of Stereotype Resources in the Sub-Saharan African Context (2026.eacl-short)

Copied to clipboard

Challenge: Existing data collection approaches to generative AI are inadequate to assess its safety and utility.
Approach: They propose a multilingual stereotype resource that uses socioculturally-situated, community-engaged methods to assess the region’s linguistic diversity and traditional orality.
Outcome: The proposed method covers four sub-Saharan African countries that are severely underrepresented in NLP resources: Ghana, Kenya, Nigeria, and South Africa.
EuroGEST: Investigating gender stereotypes in multilingual language models (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models encode social biases, but most benchmarks for gender bias remain English-centric.
Approach: They propose a dataset to measure gender-stereotypical reasoning in large language models across English and 29 European languages.
Outcome: The proposed method is highly accurate across languages and strong in translations and gender labels.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations