Getting Gender Right in Neural Machine Translation (D18-1)

Copied to clipboard

Challenge: linguistics studies show that the language used by males and females differs in terms of style and syntax.
Approach: They integrate gender information into NMT systems to improve translation quality for multiple language pairs by incorporating gender information to a large dataset.
Outcome: The proposed system significantly improves translation quality for some language pairs.

Similar Papers

Different Speech Translation Models Encode and Translate Speaker Gender Differently (2025.acl-short)

Copied to clipboard

Challenge: Recent studies on interpreting the hidden states of speech models have shown their ability to capture speaker-specific features, including gender.
Approach: They propose to use probing methods to assess gender encoding across ST models.
Outcome: The proposed models capture speaker-specific features, including gender, while older models do not . low gender encoding capabilities result in systems’ tendency toward a masculine default, a translation bias that is more pronounced in newer architectures.
Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE Corpus (2020.acl-main)

Copied to clipboard

Challenge: a growing number of studies have examined the issue of gender bias in speech translation . a gender bias is a systemic problem that reproduces gender stereotypes discriminating women.
Approach: They present the first thorough investigation of gender bias in speech translation . they compare audio technologies for English-Italian/French translations .
Outcome: The proposed method compares different technologies on two languages, English and French.
Measuring and Mitigating Name Biases in Neural Machine Translation (2022.acl-long)

Copied to clipboard

Challenge: Neural machine translation systems exhibit problematic biases, such as stereotypical gender bias in occupation terms.
Approach: They propose a method to reduce biases in person name translations by randomly switching entities during translation.
Outcome: The proposed method eliminates the problem without any effect on translation quality.
Reducing Gender Bias in Neural Machine Translation as a Domain Adaptation Problem (2020.acl-main)

Copied to clipboard

Challenge: Training data for NLP tasks often exhibits gender bias in that fewer sentences refer to women than to men.
Approach: They propose a lattice-rescoring scheme which allows a trade-off between general translation quality and bias reduction during adaptation and inference time.
Outcome: The proposed approach outperforms all systems evaluated on WinoMT with no degradation of general test set BLEU.
Investigating Failures of Automatic Translation in the Case of Unambiguous Gender (2022.acl-long)

Copied to clipboard

Challenge: Existing models are unable to make basic deductions regarding how to correctly inflect nouns with grammatical gender.
Approach: They propose to evaluate NMT models' ability to translate gender morphology correctly in unambiguous contexts across syntactically diverse sentences.
Outcome: The proposed model was unable to translate gender morphology correctly in unambiguous contexts across syntactically diverse sentences.
Integrating Language Models into Direct Speech Translation: An Inference-Time Solution to Control Gender Inflection (2023.emnlp-main)

Copied to clipboard

Challenge: Existing solutions to control speaker-related gender inflections in ST involve dedicated model retraining on gender-labeled data.
Approach: They propose to use a gender-based inference-time solution to control speaker-related gender inflections in ST by replacing the implicitly learned internal language model with gender-specific external LMs.
Outcome: The proposed approach outperforms the base models and the best training-time mitigation strategy by up to 31.0 and 1.6 points in gender accuracy, respectively, for feminine forms.
Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations (2025.findings-acl)

Copied to clipboard

Challenge: addressing gender bias and maintaining logical coherence in machine translation remains challenging, especially when translating between natural gender languages, like English, and genderless languages, such as Persian, Indonesian, and Finnish.
Approach: They propose a dataset to assess translation systems' performance in six low- to mid-resource languages and a translation dataset to examine gender bias and logical coherence.
Outcome: The Translate-with-Care dataset, comprising 3,950 challenging scenarios across six low- to mid-resource languages, reveals a universal struggle in translating genderless content, resulting in gender stereotyping and reasoning errors.
Does Context Help Mitigate Gender Bias in Neural Machine Translation? (2024.findings-emnlp)

Copied to clipboard

Challenge: Neural machine translation models perpetuate gender bias in their training data distribution.
Approach: They examine the gender bias in Neural Machine Translation by using context-aware models to enhance translation accuracy for feminine terms and translation with non-informative context in Basque to Spanish.
Outcome: The proposed models can maintain or even amplify gender bias in translations of stereotypical professions in English and with non-informative context in Basque to Spanish.
Do Multilingual Neural Machine Translation Models Contain Language Pair Specific Attention Heads? (2021.findings-acl)

Copied to clipboard

Challenge: Recent studies on multilingual representations focus on whether there is an emergence of language-independent representations or whether multilingual models partition their weights among different languages.
Approach: They analyze encoder self-attention and encoder-decoder attention heads in a multilingual neural translation model.
Outcome: The proposed model is based on a multilingual neural translation model with a language-independent representation.
Using Artificial French Data to Understand the Emergence of Gender Bias in Transformer Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies have demonstrated the ability of neural language models to learn linguistic properties without direct supervision.
Approach: They propose to use an artificial corpus generated by a PCFG to control the gender distribution in training data and determine under which conditions a model correctly captures gender information.
Outcome: The proposed approach allows to control the gender distribution in training data and determine under which conditions a model correctly captures gender information or appears gender-biased.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations