Papers by Elisa Leonardelli
Job Unfair: An Investigation of Gender and Occupational Bias in Free-Form Text Completions by LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | a recent study has identified that LLMs are used in domains where they support or replace human decision-making . a systematic review of LLM outputs shows that many facets of social bias remain unaccounted for . |
| Approach: | They propose to disentangle gender and occupational biases in Italian and English as expressed by LLMs. |
| Outcome: | The proposed method captures gender and occupational biases in Italian and English . it also shows that models struggle with gender-neutral expressions, especially beyond English - the authors conclude . |
Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators’ Disagreement (2021.emnlp-main)
Copied to clipboard
| Challenge: | supervised learning is a key component of offensive language detection, but there is little attention given to the quality of annotated data. |
| Approach: | They propose to examine the level of agreement among annotators while selecting data to create offensive language datasets, a task involving a high level of subjectivity. |
| Outcome: | The proposed datasets show that annotators' agreement has a strong effect on classifiers performance and robustness. |
Work Hard, Play Hard: Collecting Acceptability Annotations through a 3D Game (2022.lrec-1)
Copied to clipboard
| Challenge: | Corpus-based studies on acceptability judgements have always been popular thanks to the release of the CoLA corpus, a large-scale corpus of sentences extracted from linguistic handbooks as examples of acceptable/non acceptable phenomena in English. |
| Approach: | They present a 3D video game that was used to collect acceptability judgments on italian sentences and compare them with experts’ acceptability judgements. |
| Outcome: | The proposed game compares the annotations of Italian sentences with those of experts and shows that they are more reliable than crowd-sourced annotations. |
Monolingual and Cross-Lingual Acceptability Judgments with the Italian CoLA corpus (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Acceptability judgments are the most significant source of data in linguistics . however, there are still many open issues regarding methods for collecting and evaluating them. |
| Approach: | They propose to create a corpus of sentences with acceptability judgments using the same approach and the same steps as the English corpus. |
| Outcome: | The proposed corpus contains almost 10,000 sentences with acceptability judgments. |
Real Men are Tough: Evaluating Gender Bias and Sensitivity to Masculinity Norms in LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models exhibit gender bias, but most evaluations focus on downstream stereotypes . a recent study found that explicit endorsement of masculinity norms is low across models . |
| Approach: | They investigate whether large language models rely on traditional masculinity norms as latent priors in gender-biased inference. |
| Outcome: | The findings show that large language models rely on stereotypes as latent priors . the authors used the Male Role Norms Inventory (MRNI) to investigate gender bias . |
Don’t Augment, Rewrite? Assessing Abusive Language Detection with Synthetic Data (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets for abusive language detection and content moderation are limited by regulatory bodies and social media platforms. |
| Approach: | They propose to replace existing datasets in English with synthetic data by rewriting original texts with an instruction-based generative model. |
| Outcome: | The proposed model improves performance in cross-dataset training. |
Why Don’t You Do It Right? Analysing Annotators’ Disagreement in Subjective Tasks (2023.eacl-main)
Copied to clipboard
| Challenge: | Disagreement can reflect different aspects of linguistic annotation, from annotators’ subjectivity to sloppiness or lack of context to interpret a text. |
| Approach: | They propose a taxonomy of possible reasons leading to annotators' disagreement in subjective tasks and manually label part of a Twitter dataset for offensive language detection in english following this taxonomies. |
| Outcome: | The proposed taxonomy of disagreements in linguistic datasets can be used to assess how accurate tweets belonging to different disagreement categories can be classified as offensive or not. |