Multilingual Text-to-Image Generation Magnifies Gender Stereotypes (2025.acl-long)

Copied to clipboard

Challenge: Text-to-image (T2I) generation models have great results in image quality, flexibility, and text alignment, but they suffer from substantial gender bias.
Approach: They propose a benchmark to study gender bias in multilingual T2I models . they use multilingual prompts to account for grammatical differences influencing gender .
Outcome: The proposed benchmark shows strong gender biases and language-specific differences across models.

Similar Papers

Beyond Content: How Grammatical Gender Shapes Visual Representation in Text-to-Image Models (2025.findings-emnlp)

Copied to clipboard

Challenge: grammatical gender significantly influences image generation in text-to-image models . masculine grammatikal markers increase male representation to 73% on average . feminine grammatological markers increase female representation to 38% .
Approach: They propose a cross-linguistic benchmark examining words where grammatical gender contradicts stereotypical gender associations.
Outcome: The proposed benchmark examines words where grammatical gender contradicts stereotypical gender associations.
The Male CEO and the Female Assistant: Evaluation and Mitigation of Gender Biases in Text-To-Image Generation of Dual Subjects (2025.acl-long)

Copied to clipboard

Challenge: Recent large-scale T2I models like DALLE-3 have made progress in reducing gender stereotypes when generating single-person images.
Approach: They propose a framework that queries T2I models to depict two individuals with gender-stereotyped social identities to evaluate gender biases.
Outcome: The proposed framework reduces gender stereotypes when generating images with more than one person.
EuroGEST: Investigating gender stereotypes in multilingual language models (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models encode social biases, but most benchmarks for gender bias remain English-centric.
Approach: They propose a dataset to measure gender-stereotypical reasoning in large language models across English and 29 European languages.
Outcome: The proposed method is highly accurate across languages and strong in translations and gender labels.
Gender Bias in Multilingual Embeddings and Cross-Lingual Transfer (2020.acl-main)

Copied to clipboard

Challenge: Multilingual word embeddings embed words from many languages into a single semantic space such that words with similar meanings are close to each other regardless of the language.
Approach: They propose to use multilingual word embeddings to align embeddable words from multiple languages into a single semantic space so that words with similar meanings are close to each other regardless of the language.
Outcome: The proposed model can be used to learn gender bias in multilingual representations and to improve transfer learning.
Stereotypes and Smut: The (Mis)representation of Non-cisgender Identities by Text-to-Image Models (2023.findings-acl)

Copied to clipboard

Challenge: Initial studies have pointed to the potential for harm due to predictive bias, reflecting and potentially reinforcing cultural stereotypes.
Approach: They conduct a survey among non-cisgender individuals and interviews to establish which harms affected individuals anticipate, and how they would like to be represented.
Outcome: The results show that certain non-cisgender identities are consistently (mis)represented as less human, more stereotyped and more sexualised.
When Cultures Meet: Multicultural Text-to-Image Generation (2026.findings-acl)

Copied to clipboard

Challenge: a new task to evaluate text-to-image generation models for multicultural scenes is unexplored.
Approach: They propose a benchmark task to evaluate text-to-image models in multicultural settings . they use a dataset of 9,000 images spanning five countries, three age groups, two genders, 25 historical landmarks, and five languages to analyze behavior .
Outcome: The proposed benchmark analyzes the behavior of state-of-the-art models across multiple dimensions including alignment, image quality, aesthetics, knowledge, and fairness.
How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions? (2022.emnlp-main)

Copied to clipboard

Challenge: Text-to-image generative models can generate high-quality photo-realistic images conditional on natural language text descriptions in a zero-shot fashion.
Approach: They propose an Ethical NaTural Language Interventions in Text-to-Image GENeration benchmark dataset to evaluate the change in image generation conditional on ethical interventions across three social axes – gender, skin color, and culture.
Outcome: The proposed model generations cover diverse social groups while preserving image quality.
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation (2024.acl-long)

Copied to clipboard

Challenge: Existing approaches for evaluating stereotypes have a noticeable lack of coverage of global identity groups and their associated stereotypes.
Approach: They propose to use a dataset to evaluate nationality-based stereotypes in T2I models across 135 nationalities to assess offensive stereotypes.
Outcome: The proposed dataset enables evaluation of known nationality-based stereotypes across 135 nationalities.
Social Bias in Multilingual Language Models: A Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Pretrained multilingual models exhibit the same social bias as models processing English texts.
Approach: They examine the literature on bias evaluation and mitigation approaches in multilingual and non-English contexts and identify gaps in the field.
Outcome: The proposed models perform well on multilingual language understanding benchmarks and are consistent with the current literature.
On Evaluating and Mitigating Gender Biases in Multilingual Settings (2023.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks and resources for evaluating gender biases in multilingual settings are limited.
Approach: They propose to extend DisCo to different Indian languages using human annotations to evaluate gender biases in multilingual models.
Outcome: The proposed benchmarks and mitigation techniques are extended beyond English to evaluate gender biases in multilingual models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations