Papers by Elisa Leonardelli

7 papers
Job Unfair: An Investigation of Gender and Occupational Bias in Free-Form Text Completions by LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: a recent study has identified that LLMs are used in domains where they support or replace human decision-making . a systematic review of LLM outputs shows that many facets of social bias remain unaccounted for .
Approach: They propose to disentangle gender and occupational biases in Italian and English as expressed by LLMs.
Outcome: The proposed method captures gender and occupational biases in Italian and English . it also shows that models struggle with gender-neutral expressions, especially beyond English - the authors conclude .
Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators’ Disagreement (2021.emnlp-main)

Copied to clipboard

Challenge: supervised learning is a key component of offensive language detection, but there is little attention given to the quality of annotated data.
Approach: They propose to examine the level of agreement among annotators while selecting data to create offensive language datasets, a task involving a high level of subjectivity.
Outcome: The proposed datasets show that annotators' agreement has a strong effect on classifiers performance and robustness.
Work Hard, Play Hard: Collecting Acceptability Annotations through a 3D Game (2022.lrec-1)

Copied to clipboard

Challenge: Corpus-based studies on acceptability judgements have always been popular thanks to the release of the CoLA corpus, a large-scale corpus of sentences extracted from linguistic handbooks as examples of acceptable/non acceptable phenomena in English.
Approach: They present a 3D video game that was used to collect acceptability judgments on italian sentences and compare them with experts’ acceptability judgements.
Outcome: The proposed game compares the annotations of Italian sentences with those of experts and shows that they are more reliable than crowd-sourced annotations.
Monolingual and Cross-Lingual Acceptability Judgments with the Italian CoLA corpus (2021.findings-emnlp)

Copied to clipboard

Challenge: Acceptability judgments are the most significant source of data in linguistics . however, there are still many open issues regarding methods for collecting and evaluating them.
Approach: They propose to create a corpus of sentences with acceptability judgments using the same approach and the same steps as the English corpus.
Outcome: The proposed corpus contains almost 10,000 sentences with acceptability judgments.
Real Men are Tough: Evaluating Gender Bias and Sensitivity to Masculinity Norms in LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Large language models exhibit gender bias, but most evaluations focus on downstream stereotypes . a recent study found that explicit endorsement of masculinity norms is low across models .
Approach: They investigate whether large language models rely on traditional masculinity norms as latent priors in gender-biased inference.
Outcome: The findings show that large language models rely on stereotypes as latent priors . the authors used the Male Role Norms Inventory (MRNI) to investigate gender bias .
Don’t Augment, Rewrite? Assessing Abusive Language Detection with Synthetic Data (2024.findings-acl)

Copied to clipboard

Challenge: Existing datasets for abusive language detection and content moderation are limited by regulatory bodies and social media platforms.
Approach: They propose to replace existing datasets in English with synthetic data by rewriting original texts with an instruction-based generative model.
Outcome: The proposed model improves performance in cross-dataset training.
Why Don’t You Do It Right? Analysing Annotators’ Disagreement in Subjective Tasks (2023.eacl-main)

Copied to clipboard

Challenge: Disagreement can reflect different aspects of linguistic annotation, from annotators’ subjectivity to sloppiness or lack of context to interpret a text.
Approach: They propose a taxonomy of possible reasons leading to annotators' disagreement in subjective tasks and manually label part of a Twitter dataset for offensive language detection in english following this taxonomies.
Outcome: The proposed taxonomy of disagreements in linguistic datasets can be used to assess how accurate tweets belonging to different disagreement categories can be classified as offensive or not.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations