Papers by Debora Nozza
Countering Hateful and Offensive Speech Online - Open Challenges (2024.emnlp-tutorials)
Copied to clipboard
Leon Derczynski, Marco Guerini, Debora Nozza, Flor Miriam Plaza-del-Arco, Jeffrey Sorensen, Marcos Zampieri
| Challenge: | a comprehensive understanding of the field is needed to maintain respectful and inclusive online environments. |
| Approach: | This tutorial aims to provide attendees with a comprehensive understanding of the field by delving into essential dimensions such as multilingualism, counter-narrative generation, a hands-on session with one of the most popular APIs for detecting hate speech, fairness, and ethics in AI, and the use of recent advanced approaches. |
| Outcome: | This tutorial aims to provide attendees with a comprehensive understanding of the field by delving into essential dimensions such as multilingualism, counter-narrative generation, a hands-on session with one of the most popular APIs for detecting hate speech, fairness, and ethics in AI, and the use of recent advanced approaches. |
Personalization up to a Point: Why Personalized Content Moderation Needs Boundaries, and How We Can Enforce Them (2025.emnlp-main)
Copied to clipboard
| Challenge: | Personalized content moderation can protect users from harm while facilitating free expression . however, it can also allow highly harmful and even illegal hate speech to spread . |
| Approach: | They propose to enforce legal boundaries on personalized content moderation models to reduce legal violations while maintaining user welfare. |
| Outcome: | The proposed approach reduces legal violations while maintaining user welfare while maintaining a high degree of model performance. |
A Tale of Pronouns: Interpretability Informs Gender Bias Mitigation for Fairer Instruction-Tuned Machine Translation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent instruction fine-tuned models can solve multiple NLP tasks when prompted to do so, but current research focuses on standard performance benchmarks, leaving compelling fairness and ethical considerations behind. |
| Approach: | They propose a method to mitigate gender bias in machine translation by using a corpus of machine translations from the WinoMT corpus. |
| Outcome: | The proposed model can solve multiple NLP tasks when prompted, but it lacks fairness and ethical considerations. |
PATS: Personality-Aware Teaching Strategies with Large Language Model Tutors (2026.findings-eacl)
Copied to clipboard
Donya Rooein, Sankalan Pal Chowdhury, Mariia Eremeeva, Yuan Qin, Debora Nozza, Mrinmaya Sachan, Dirk Hovy
| Challenge: | pedagogical theories are not aligned with teaching strategies for educational tasks . quiet students may be disengaged or not thinking critically because they do not speak up . |
| Approach: | They propose a taxonomy that links pedagogical methods to personality profiles to map teaching strategies to student personality traits. |
| Outcome: | The proposed model improves the use of less common, high-impact strategies such as role-playing . the model also increases the use less common strategies such role-players . |
ferret: a Framework for Benchmarking Explainers on Transformers (2023.eacl-demo)
Copied to clipboard
| Challenge: | Existing methods for interpreting transformer outputs are scattered and hard to operationalize. |
| Approach: | They propose a Python library to simplify the use and comparisons of XAI methods on transformers. |
| Outcome: | The proposed method provides better explanations and is preferable in the context of transformer models. |
The State of Profanity Obfuscation in Natural Language Processing Scientific Publications (2023.findings-acl)
Copied to clipboard
| Challenge: | obfuscation is used for English but not other languages, and even then, unevenly. |
| Approach: | They propose a multilingual community resource called PrOf to standardize profanity obfuscation processes. |
| Outcome: | The proposed tool can help scientific publications to make hate speech work accessible and comparable, irrespective of language. |
Exposing the limits of Zero-shot Cross-lingual Hate Speech Detection (2021.acl-short)
Copied to clipboard
| Challenge: | a lack of labeled, non-English resources for hate speech detection limits research on hate speech . a recent study shows that zero-shot, cross-lingual learning models cannot be used as they are . lack of consistency limits research, and lack of models for non-english languages limits learning . |
| Approach: | They propose a zero-shot, cross-lingual transfer learning framework for hate speech detection . they use benchmark data sets in English, Italian, and Spanish to detect hate speech . |
| Outcome: | The proposed framework can't be used as it is, but needs to be carefully designed, the authors say . they find that non-hateful, language-specific taboo interjections are misinterpreted as signals of hate speech . |
What about “em”? How Commercial Machine Translation Fails to Handle (Neo-)Pronouns (2023.acl-long)
Copied to clipboard
| Challenge: | Wrong pronoun translations can discriminate against marginalized groups, e.g., non-binary individuals. |
| Approach: | They compare 3rd-person pronoun translations to five other languages . they propose to address gender exclusivity in future research . |
| Outcome: | The proposed method compares translations of gendered vs. gender-neutral pronouns from english to five other languages and vice versa. |
Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLP (2024.emnlp-main)
Copied to clipboard
| Challenge: | a measure’s intended use and reliability assessment are often unclear or entirely absent from the literature examining bias measures in natural language processing. |
| Approach: | They propose a set of desiderata to assess the degree to which a measure’s results enable informed action and a review of 146 papers proposing bias measures in NLP. |
| Outcome: | The proposed desiderata are based on 146 papers proposing bias measures in natural language processing (NLP) . they show that key elements of actionability, including a measure’s intended use and reliability assessment, are often unclear or entirely absent. |
HONEST: Measuring Hurtful Sentence Completion in Language Models (2021.naacl-main)
Copied to clipboard
| Challenge: | 4.3% of the time, language models complete a sentence with a hurtful word . authors propose a score to quantify the amount of hurtful sentence completions in a language model. |
| Approach: | They propose a score to measure hurtful sentence completions in language models . they use a template- and lexicon-based bias evaluation methodology for six languages . |
| Outcome: | The proposed score measures the amount of hurtful sentences in language models. |
The Pluralistic Moral Gap: Understanding Moral Judgment and Value Differences between Humans and Large Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing studies have shown that Large Language Models (LLMs) are not fully aligned with human moral judgments. |
| Approach: | They propose a dataset of 1,618 real-world moral dilemmas paired with a distribution of human moral judgments consisting of a binary evaluation and a free-text rationale to examine how closely LLMs align with human moral judgements. |
| Outcome: | The proposed model reproduces human judgments only under high consensus; alignment deteriorates sharply when human disagreement increases. |
Cross-lingual Contextualized Topic Models with Zero-shot Learning (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing topic models are language-specific and cannot be transferred in a transferable manner. |
| Approach: | They propose a zero-shot cross-lingual topic model that learns topics on one language and predicts them for unseen documents in different languages. |
| Outcome: | The proposed model learns topics on one language and predicts them for unseen documents in different languages. |
Can I Introduce My Boyfriend to My Grandmother? Evaluating Large Language Models Capabilities on Iranian Social Norm Classification (2025.findings-naacl)
Copied to clipboard
| Challenge: | Introducing the Iranian Social Norms dataset, a collection of 1,699 social norms, with Farsi adding linguistic complexity. |
| Approach: | They propose a collection of Iranian social norms with English translations and a novel Iranian dataset. |
| Outcome: | The Iranian Social Norms dataset is the first to be used in the Farsi language . it includes 1,699 social norms including environments, demographic features, and scope annotation, alongside English translations. |
The “r” in “woman” stands for rights. Auditing LLMs in Uncovering Social Dynamics in Implicit Misogyny (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a recent study examined misogynistic expressions in English and Italian . a taxonomy of social dynamics is used to identify misogorical expressions . |
| Approach: | They examine misogynistic expressions in English and Italian using a taxonomy of social dynamics . they find that LLMs struggle to follow instructions and reason in all settings . |
| Outcome: | The results show that misogynistic expressions are more often implicit than openly hostile . the authors show that LLMs struggle to follow instructions and reason in all settings . |
A Multi-dimensional study on Bias in Vision-Language models (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent studies have focused on the issue of bias in joint Vision-Language models . pre-trained models complete a neutral template with a hurtful word 5% of the time . |
| Approach: | They propose to use a multi-dimensional bias metric to investigate bias in English VL models . they use gender, ethnicity, and age as dimensions to analyze bias in VLs . |
| Outcome: | The proposed model is based on gender, ethnicity, and age as dimensions. |
Biased Tales: Cultural and Topic Bias in Generating Children’s Stories (2025.emnlp-main)
Copied to clipboard
| Challenge: | Personalized stories are often preferred because they reflect a child's interests, experiences, and developmental needs. |
| Approach: | They analyze a dataset to examine how biases influence protagonists’ attributes and story elements in LLM-generated stories. |
| Outcome: | The proposed dataset shows that gender stereotypes influence protagonist attributes and story elements in LLM-generated stories. |
Entropy-based Attention Regularization Frees Unintended Bias Mitigation from Lists (2022.findings-acl)
Copied to clipboard
| Challenge: | E.g., neural hate speech detection models are strongly influenced by identity terms like gay, or women, resulting in false positives, severe unintended bias, and lower performance. |
| Approach: | They propose a knowledge-free Entropy-based Attention Regularization (EAR) approach to discourage overfitting to training-specific terms. |
| Outcome: | The proposed model matches or exceeds state-of-the-art performance for hate speech classification and bias metrics on three benchmark corpora in English and Italian. |
Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages (2022.emnlp-main)
Copied to clipboard
| Challenge: | Hate speech datasets focus on English-language content, hindering effective models . annotating hateful content is expensive, time-consuming and potentially harmful to annotators. |
| Approach: | They propose to use ISO 639-1 codes to fine-tune models on one source language and apply them to another language. |
| Outcome: | The proposed approach performs well on some tasks, but fails on many others. |