Papers by Debora Nozza

18 papers
Countering Hateful and Offensive Speech Online - Open Challenges (2024.emnlp-tutorials)

Copied to clipboard

Challenge: a comprehensive understanding of the field is needed to maintain respectful and inclusive online environments.
Approach: This tutorial aims to provide attendees with a comprehensive understanding of the field by delving into essential dimensions such as multilingualism, counter-narrative generation, a hands-on session with one of the most popular APIs for detecting hate speech, fairness, and ethics in AI, and the use of recent advanced approaches.
Outcome: This tutorial aims to provide attendees with a comprehensive understanding of the field by delving into essential dimensions such as multilingualism, counter-narrative generation, a hands-on session with one of the most popular APIs for detecting hate speech, fairness, and ethics in AI, and the use of recent advanced approaches.
Personalization up to a Point: Why Personalized Content Moderation Needs Boundaries, and How We Can Enforce Them (2025.emnlp-main)

Copied to clipboard

Challenge: Personalized content moderation can protect users from harm while facilitating free expression . however, it can also allow highly harmful and even illegal hate speech to spread .
Approach: They propose to enforce legal boundaries on personalized content moderation models to reduce legal violations while maintaining user welfare.
Outcome: The proposed approach reduces legal violations while maintaining user welfare while maintaining a high degree of model performance.
A Tale of Pronouns: Interpretability Informs Gender Bias Mitigation for Fairer Instruction-Tuned Machine Translation (2023.emnlp-main)

Copied to clipboard

Challenge: Recent instruction fine-tuned models can solve multiple NLP tasks when prompted to do so, but current research focuses on standard performance benchmarks, leaving compelling fairness and ethical considerations behind.
Approach: They propose a method to mitigate gender bias in machine translation by using a corpus of machine translations from the WinoMT corpus.
Outcome: The proposed model can solve multiple NLP tasks when prompted, but it lacks fairness and ethical considerations.
PATS: Personality-Aware Teaching Strategies with Large Language Model Tutors (2026.findings-eacl)

Copied to clipboard

Challenge: pedagogical theories are not aligned with teaching strategies for educational tasks . quiet students may be disengaged or not thinking critically because they do not speak up .
Approach: They propose a taxonomy that links pedagogical methods to personality profiles to map teaching strategies to student personality traits.
Outcome: The proposed model improves the use of less common, high-impact strategies such as role-playing . the model also increases the use less common strategies such role-players .
ferret: a Framework for Benchmarking Explainers on Transformers (2023.eacl-demo)

Copied to clipboard

Challenge: Existing methods for interpreting transformer outputs are scattered and hard to operationalize.
Approach: They propose a Python library to simplify the use and comparisons of XAI methods on transformers.
Outcome: The proposed method provides better explanations and is preferable in the context of transformer models.
The State of Profanity Obfuscation in Natural Language Processing Scientific Publications (2023.findings-acl)

Copied to clipboard

Challenge: obfuscation is used for English but not other languages, and even then, unevenly.
Approach: They propose a multilingual community resource called PrOf to standardize profanity obfuscation processes.
Outcome: The proposed tool can help scientific publications to make hate speech work accessible and comparable, irrespective of language.
Exposing the limits of Zero-shot Cross-lingual Hate Speech Detection (2021.acl-short)

Copied to clipboard

Challenge: a lack of labeled, non-English resources for hate speech detection limits research on hate speech . a recent study shows that zero-shot, cross-lingual learning models cannot be used as they are . lack of consistency limits research, and lack of models for non-english languages limits learning .
Approach: They propose a zero-shot, cross-lingual transfer learning framework for hate speech detection . they use benchmark data sets in English, Italian, and Spanish to detect hate speech .
Outcome: The proposed framework can't be used as it is, but needs to be carefully designed, the authors say . they find that non-hateful, language-specific taboo interjections are misinterpreted as signals of hate speech .
What about “em”? How Commercial Machine Translation Fails to Handle (Neo-)Pronouns (2023.acl-long)

Copied to clipboard

Challenge: Wrong pronoun translations can discriminate against marginalized groups, e.g., non-binary individuals.
Approach: They compare 3rd-person pronoun translations to five other languages . they propose to address gender exclusivity in future research .
Outcome: The proposed method compares translations of gendered vs. gender-neutral pronouns from english to five other languages and vice versa.
Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLP (2024.emnlp-main)

Copied to clipboard

Challenge: a measure’s intended use and reliability assessment are often unclear or entirely absent from the literature examining bias measures in natural language processing.
Approach: They propose a set of desiderata to assess the degree to which a measure’s results enable informed action and a review of 146 papers proposing bias measures in NLP.
Outcome: The proposed desiderata are based on 146 papers proposing bias measures in natural language processing (NLP) . they show that key elements of actionability, including a measure’s intended use and reliability assessment, are often unclear or entirely absent.
HONEST: Measuring Hurtful Sentence Completion in Language Models (2021.naacl-main)

Copied to clipboard

Challenge: 4.3% of the time, language models complete a sentence with a hurtful word . authors propose a score to quantify the amount of hurtful sentence completions in a language model.
Approach: They propose a score to measure hurtful sentence completions in language models . they use a template- and lexicon-based bias evaluation methodology for six languages .
Outcome: The proposed score measures the amount of hurtful sentences in language models.
The Pluralistic Moral Gap: Understanding Moral Judgment and Value Differences between Humans and Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies have shown that Large Language Models (LLMs) are not fully aligned with human moral judgments.
Approach: They propose a dataset of 1,618 real-world moral dilemmas paired with a distribution of human moral judgments consisting of a binary evaluation and a free-text rationale to examine how closely LLMs align with human moral judgements.
Outcome: The proposed model reproduces human judgments only under high consensus; alignment deteriorates sharply when human disagreement increases.
Cross-lingual Contextualized Topic Models with Zero-shot Learning (2021.eacl-main)

Copied to clipboard

Challenge: Existing topic models are language-specific and cannot be transferred in a transferable manner.
Approach: They propose a zero-shot cross-lingual topic model that learns topics on one language and predicts them for unseen documents in different languages.
Outcome: The proposed model learns topics on one language and predicts them for unseen documents in different languages.
Can I Introduce My Boyfriend to My Grandmother? Evaluating Large Language Models Capabilities on Iranian Social Norm Classification (2025.findings-naacl)

Copied to clipboard

Challenge: Introducing the Iranian Social Norms dataset, a collection of 1,699 social norms, with Farsi adding linguistic complexity.
Approach: They propose a collection of Iranian social norms with English translations and a novel Iranian dataset.
Outcome: The Iranian Social Norms dataset is the first to be used in the Farsi language . it includes 1,699 social norms including environments, demographic features, and scope annotation, alongside English translations.
The “r” in “woman” stands for rights. Auditing LLMs in Uncovering Social Dynamics in Implicit Misogyny (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study examined misogynistic expressions in English and Italian . a taxonomy of social dynamics is used to identify misogorical expressions .
Approach: They examine misogynistic expressions in English and Italian using a taxonomy of social dynamics . they find that LLMs struggle to follow instructions and reason in all settings .
Outcome: The results show that misogynistic expressions are more often implicit than openly hostile . the authors show that LLMs struggle to follow instructions and reason in all settings .
A Multi-dimensional study on Bias in Vision-Language models (2023.findings-acl)

Copied to clipboard

Challenge: Recent studies have focused on the issue of bias in joint Vision-Language models . pre-trained models complete a neutral template with a hurtful word 5% of the time .
Approach: They propose to use a multi-dimensional bias metric to investigate bias in English VL models . they use gender, ethnicity, and age as dimensions to analyze bias in VLs .
Outcome: The proposed model is based on gender, ethnicity, and age as dimensions.
Biased Tales: Cultural and Topic Bias in Generating Children’s Stories (2025.emnlp-main)

Copied to clipboard

Challenge: Personalized stories are often preferred because they reflect a child's interests, experiences, and developmental needs.
Approach: They analyze a dataset to examine how biases influence protagonists’ attributes and story elements in LLM-generated stories.
Outcome: The proposed dataset shows that gender stereotypes influence protagonist attributes and story elements in LLM-generated stories.
Entropy-based Attention Regularization Frees Unintended Bias Mitigation from Lists (2022.findings-acl)

Copied to clipboard

Challenge: E.g., neural hate speech detection models are strongly influenced by identity terms like gay, or women, resulting in false positives, severe unintended bias, and lower performance.
Approach: They propose a knowledge-free Entropy-based Attention Regularization (EAR) approach to discourage overfitting to training-specific terms.
Outcome: The proposed model matches or exceeds state-of-the-art performance for hate speech classification and bias metrics on three benchmark corpora in English and Italian.
Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages (2022.emnlp-main)

Copied to clipboard

Challenge: Hate speech datasets focus on English-language content, hindering effective models . annotating hateful content is expensive, time-consuming and potentially harmful to annotators.
Approach: They propose to use ISO 639-1 codes to fine-tune models on one source language and apply them to another language.
Outcome: The proposed approach performs well on some tasks, but fails on many others.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations