Papers by Giuseppe Attanasio

14 papers
Building Bridges: A Dataset for Evaluating Gender-Fair Machine Translation into German (2024.findings-acl)

Copied to clipboard

Challenge: a societal movement towards using gender-fair language exists, but gender-free German is barely supported in machine translation.
Approach: They propose to use a community-created gender-fair language dictionary to study gender-neutral German . they also use encyclopedic text and parliamentary speeches to translate the words in isolation .
Outcome: The proposed study shows that most systems produce mainly masculine forms and rarely gender-neutral variants.
A Tale of Pronouns: Interpretability Informs Gender Bias Mitigation for Fairer Instruction-Tuned Machine Translation (2023.emnlp-main)

Copied to clipboard

Challenge: Recent instruction fine-tuned models can solve multiple NLP tasks when prompted to do so, but current research focuses on standard performance benchmarks, leaving compelling fairness and ethical considerations behind.
Approach: They propose a method to mitigate gender bias in machine translation by using a corpus of machine translations from the WinoMT corpus.
Outcome: The proposed model can solve multiple NLP tasks when prompted, but it lacks fairness and ethical considerations.
Different Speech Translation Models Encode and Translate Speaker Gender Differently (2025.acl-short)

Copied to clipboard

Challenge: Recent studies on interpreting the hidden states of speech models have shown their ability to capture speaker-specific features, including gender.
Approach: They propose to use probing methods to assess gender encoding across ST models.
Outcome: The proposed models capture speaker-specific features, including gender, while older models do not . low gender encoding capabilities result in systems’ tendency toward a masculine default, a translation bias that is more pronounced in newer architectures.
Classist Tools: Social Class Correlates with Performance in NLP (2024.acl-long)

Copied to clipboard

Challenge: despite growing concerns surrounding fairness and bias in NLP, there is a dearth of studies delving into the effects it may have on NLP systems.
Approach: They argue that NLP systems’ performance is affected by speakers’ SES, potentially disadvantaging less-privileged socioeconomic groups.
Outcome: The proposed model shows that NLP systems perform better on tasks with social class, ethnicity and geographical variation than those without social class.
ferret: a Framework for Benchmarking Explainers on Transformers (2023.eacl-demo)

Copied to clipboard

Challenge: Existing methods for interpreting transformer outputs are scattered and hard to operationalize.
Approach: They propose a Python library to simplify the use and comparisons of XAI methods on transformers.
Outcome: The proposed method provides better explanations and is preferable in the context of transformer models.
Explaining Speech Classification Models via Word-Level Audio Segments and Paralinguistic Features (2024.eacl-long)

Copied to clipboard

Challenge: Existing explanations for speech classification models are difficult to interpret and make mistakes.
Approach: They propose to explain speech classification models by using word-level and paralinguistic attributes to measure the impact of each audio segment aligned with a word on the outcome.
Outcome: The proposed explanations correctly represent the model’s inner workings and are plausible to humans.
Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation (2025.acl-long)

Copied to clipboard

Challenge: Qualitative estimation (QE) metrics have been optimized to align with human quality judgments, but whether they encode social biases has been largely overlooked.
Approach: They define and investigate gender bias of QE metrics and discuss its downstream implications for machine translation (MT) when a human entity’s gender in the source is undisclosed, masculine-inflected translations score higher than feminine-infflectes translations are penalized.
Outcome: The proposed measures are based on gender-based quality estimation metrics across multiple domains, datasets, and languages.
Twists, Humps, and Pebbles: Multilingual Speech Recognition Models Exhibit Gender Performance Gaps (2024.emnlp-main)

Copied to clipboard

Challenge: a new class of multitasks, multilingual neural networks, has recently pushed the boundaries of speech-related tasks.
Approach: They evaluate performance of two widely used multilingual automatic speech recognition models . they find clear gender disparities, with the advantaged group varying across languages .
Outcome: The proposed models are compared on 19 languages from eight language families and two speaking conditions.
Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLP (2024.emnlp-main)

Copied to clipboard

Challenge: a measure’s intended use and reliability assessment are often unclear or entirely absent from the literature examining bias measures in natural language processing.
Approach: They propose a set of desiderata to assess the degree to which a measure’s results enable informed action and a review of 146 papers proposing bias measures in NLP.
Outcome: The proposed desiderata are based on 146 papers proposing bias measures in natural language processing (NLP) . they show that key elements of actionability, including a measure’s intended use and reliability assessment, are often unclear or entirely absent.
Glitter: A Multi-Sentence, Multi-Reference Benchmark for Gender-Fair German Machine Translation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing MT models are limited in size and often consist of single sentences or single gender-fair formulation types.
Approach: They propose a benchmark for machine translation that features extended passages with professional translations implementing gender-fair alternatives: neutral rewording, typographical solutions and neologistic forms.
Outcome: The proposed benchmark features extended passages with professional translations implementing three gender-fair alternatives: neutral rewording, typographical solutions (gender star), and neologistic forms (-ens forms).
Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTE (2025.emnlp-main)

Copied to clipboard

Challenge: Genderneutral translation (GNT) is a linguistic strategy towards fairer communication across languages.
Approach: They propose to use a multilingual evaluation resource to evaluate inclusive translation with state-of-the-art instruction-following language models (LMs)
Outcome: The proposed model can recognize when neutrality is appropriate, but cannot consistently produce neutral translations, limiting their usability.
Entropy-based Attention Regularization Frees Unintended Bias Mitigation from Lists (2022.findings-acl)

Copied to clipboard

Challenge: E.g., neural hate speech detection models are strongly influenced by identity terms like gay, or women, resulting in false positives, severe unintended bias, and lower performance.
Approach: They propose a knowledge-free Entropy-based Attention Regularization (EAR) approach to discourage overfitting to training-specific terms.
Outcome: The proposed model matches or exceeds state-of-the-art performance for hate speech classification and bias metrics on three benchmark corpora in English and Italian.
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are now being used by millions of people across the world.
Approach: They propose a test suite called XSTest to identify such eXaggerated Safety behaviours in a systematic way.
Outcome: The proposed test suite identifies eXaggerated Safety behaviours in a systematic way.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations