Papers by Simona Frenda

9 papers
POLAR: A Benchmark for Multilingual, Multicultural, and Multi-Event Online Polarization (2026.findings-acl)

Copied to clipboard

Challenge: polarization is a pervasive threat to democratic institutions, civil discourse, and social cohesion worldwide . most existing datasets focus on English or high-resource languages, reflecting a widespread trend across NLP tasks .
Approach: They propose a multilingual, multicultural, and multi-event dataset with over 110K instances in 22 languages drawn from diverse online platforms and real-world events.
Outcome: The proposed dataset analyzes polarization detection, type, and manifestation using a variety of annotation platforms adapted to each cultural context.
Human vs. Machine Perceptions on Immigration Stereotypes (2024.lrec-main)

Copied to clipboard

Challenge: a growing number of natural language processing models leave aside the language itself . a recent paradigm in the computational linguistics community is training models on specific perspectives of a segment of the population or an individual.
Approach: They propose to use BERT-based classification models to detect stereotypes related to immigrants . they compare models with predictions from GPT-4 and annotated tweets from Spanish Twitter .
Outcome: The proposed models are compared with predictions from the dataset of Spanish Twitter posts containing stereotypes . the models are confident in their predictions and more accurate for implicit stereotypes, the authors show .
QUEEREOTYPES: A Multi-Source Italian Corpus of Stereotypes towards LGBTQIA+ Community Members (2024.lrec-main)

Copied to clipboard

Challenge: a dataset of social media texts addressing LGBTQIA+ individuals is presented in this paper . the dataset is based on two sources in italian: Facebook and Twitter .
Approach: They describe a dataset composed of two sub-corpora from two different sources in Italian . the dataset includes social media texts regarding LGBTQIA+ individuals, behaviors, ideology and events .
Outcome: The QUEEREOTYPES dataset includes social media texts regarding LGBTQIA+ individuals, behaviors, ideology and events.
EPIC: Multi-Perspective Annotation of a Corpus of Irony (2023.acl-long)

Copied to clipboard

Challenge: EPIC is the first annotated corpus for irony analysis based on data perspectivism . a recent trend in natural language processing (NLP) postulates that the disagreement among annotators in a language resource is a valuable source of knowledge, rather than noise that ought to be minimized or discarded.
Approach: They propose to annotate an English perspectivist irony corpus based on data perspectivism . they validate the model by creating perspective-aware models that encode the perspectives of annotators grouped according to their demographic characteristics.
Outcome: The proposed model can capture different perspectives on irony among different groups of annotators, and is more confident than non-perspectivist models.
Confidence-based Ensembling of Perspective-aware Models (2023.emnlp-main)

Copied to clipboard

Challenge: Human label variability has been a topic of research in the field of NLP recently . Exploiting disagreements in annotations has been shown to offer advantages for accurate modelling and fairer evaluation.
Approach: They propose a highly perspectivist model that exploits disagreements in annotations to capture the subjectivity encoded in the annotation process.
Outcome: The proposed model is validated on irony and hate speech detection scenarios in in-domain and cross-domain settings.
Counterspeech Generation using Small Language Models (2026.acl-srw)

Copied to clipboard

Challenge: Social media use is growing annually with about 68.5% of the global population active on these platforms as of July 2025.
Approach: They evaluate SLMs ranging from 100 million to 3 billion parameters using simple prompting strategies as well as fine-tuning, combining automatic and robust human evaluations.
Outcome: The proposed models generate relevant, coherent, and high-quality counterspeech, suggesting their suitability for efficient and responsible deployments.
A Multilingual Dataset of Racial Stereotypes in Social Media Conversational Threads (2023.findings-eacl)

Copied to clipboard

Challenge: a new corpus-based study addresses racial stereotypes in social media conversations . a multilingual corpus of rhs is used to investigate how they are spread .
Approach: They propose a corpus-based method for multilingual racial stereotype identification in social media conversational threads.
Outcome: The proposed method sheds light on how racial hoaxes are spread and allows identification of negative stereotypes that reinforce them.
Are you sure? Measuring models bias in content moderation through uncertainty (2025.findings-emnlp)

Copied to clipboard

Challenge: Language Model-based classifiers perpetuate racial and social biases in content moderation . et al., j. n. d., and j neil, e. c. (2005) measure the fairness of content moderated models .
Approach: They propose an unsupervised approach that benchmarks models on their uncertainty . they use uncertainty as a proxy to analyze the bias of 11 models against women and non-whites .
Outcome: The proposed method analyzes the bias of 11 models against women and non-white annotators . it shows that some pre-trained models predict with high accuracy the labels coming from minority groups .
APPReddit: a Corpus of Reddit Posts Annotated for Appraisal (2022.lrec-1)

Copied to clipboard

Challenge: Existing resources for emotion recognition are lacking for appraisal models.
Approach: They propose to use APPReddit to annotate non-experimental data according to Appraisal theories . they compare it with enISEAR, a corpus of events created in an experimental setting and annotated according to this theory.
Outcome: The proposed model predicts four appraisal dimensions without significant loss . the proposed model is compared with enISEAR, a corpus of events created in an experimental setting and annotated for appraisal.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations