Papers by Camilla Casula

6 papers
Job Unfair: An Investigation of Gender and Occupational Bias in Free-Form Text Completions by LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: a recent study has identified that LLMs are used in domains where they support or replace human decision-making . a systematic review of LLM outputs shows that many facets of social bias remain unaccounted for .
Approach: They propose to disentangle gender and occupational biases in Italian and English as expressed by LLMs.
Outcome: The proposed method captures gender and occupational biases in Italian and English . it also shows that models struggle with gender-neutral expressions, especially beyond English - the authors conclude .
Generation-Based Data Augmentation for Offensive Language Detection: Is It Worth It? (2023.eacl-main)

Copied to clipboard

Challenge: generative data augmentation has been shown to be effective in offensive language detection but the potential for bias injection has not been investigated.
Approach: They propose to investigate the robustness of models trained on generated data in a variety of data augmentation setups and analyze models using the HateCheck suite.
Outcome: The proposed model training setups on four English offensive language datasets are robust and robust, while the generative DA setups do not present bias injection issues.
Real Men are Tough: Evaluating Gender Bias and Sensitivity to Masculinity Norms in LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Large language models exhibit gender bias, but most evaluations focus on downstream stereotypes . a recent study found that explicit endorsement of masculinity norms is low across models .
Approach: They investigate whether large language models rely on traditional masculinity norms as latent priors in gender-biased inference.
Outcome: The findings show that large language models rely on stereotypes as latent priors . the authors used the Male Role Norms Inventory (MRNI) to investigate gender bias .
Don’t Augment, Rewrite? Assessing Abusive Language Detection with Synthetic Data (2024.findings-acl)

Copied to clipboard

Challenge: Existing datasets for abusive language detection and content moderation are limited by regulatory bodies and social media platforms.
Approach: They propose to replace existing datasets in English with synthetic data by rewriting original texts with an instruction-based generative model.
Outcome: The proposed model improves performance in cross-dataset training.
Variationist: Exploring Multifaceted Variation and Bias in Written Language Data (2024.acl-demos)

Copied to clipboard

Challenge: Existing tools that inspect and visualize language data are limited in their capabilities.
Approach: They propose a highly-modular, extensible, and task-agnostic tool that inspects language variation and bias across multiple variables, language units, and diverse metrics.
Outcome: The proposed tool can inspect and visualize language variation and bias across variables, language units, and diverse metrics that go beyond descriptive statistics.
Delving into Qualitative Implications of Synthetic Data for Hate Speech Detection (2024.emnlp-main)

Copied to clipboard

Challenge: Recent work on synthetic data for training models for NLP tasks reports mixed results on subjective tasks such as hate speech detection.
Approach: They propose to use synthetic data to train models for highly subjective tasks such as hate speech detection to investigate the potential and specific pitfalls of using it.
Outcome: The proposed model outperforms models trained with real data on hate speech detection tasks, but it fails to accurately reflect real-world data on linguistic dimensions and results in different class distributions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations