Multilingual Text-to-Image Generation Magnifies Gender Stereotypes (2025.acl-long)
Copied to clipboard
Felix Friedrich, Katharina Hämmerl, Patrick Schramowski, Manuel Brack, Jindřich Libovický, Alexander Fraser, Kristian Kersting
| Challenge: | Text-to-image (T2I) generation models have great results in image quality, flexibility, and text alignment, but they suffer from substantial gender bias. |
| Approach: | They propose a benchmark to study gender bias in multilingual T2I models . they use multilingual prompts to account for grammatical differences influencing gender . |
| Outcome: | The proposed benchmark shows strong gender biases and language-specific differences across models. |
Similar Papers
Beyond Content: How Grammatical Gender Shapes Visual Representation in Text-to-Image Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | grammatical gender significantly influences image generation in text-to-image models . masculine grammatikal markers increase male representation to 73% on average . feminine grammatological markers increase female representation to 38% . |
| Approach: | They propose a cross-linguistic benchmark examining words where grammatical gender contradicts stereotypical gender associations. |
| Outcome: | The proposed benchmark examines words where grammatical gender contradicts stereotypical gender associations. |
The Male CEO and the Female Assistant: Evaluation and Mitigation of Gender Biases in Text-To-Image Generation of Dual Subjects (2025.acl-long)
Copied to clipboard
| Challenge: | Recent large-scale T2I models like DALLE-3 have made progress in reducing gender stereotypes when generating single-person images. |
| Approach: | They propose a framework that queries T2I models to depict two individuals with gender-stereotyped social identities to evaluate gender biases. |
| Outcome: | The proposed framework reduces gender stereotypes when generating images with more than one person. |
EuroGEST: Investigating gender stereotypes in multilingual language models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models encode social biases, but most benchmarks for gender bias remain English-centric. |
| Approach: | They propose a dataset to measure gender-stereotypical reasoning in large language models across English and 29 European languages. |
| Outcome: | The proposed method is highly accurate across languages and strong in translations and gender labels. |
Gender Bias in Multilingual Embeddings and Cross-Lingual Transfer (2020.acl-main)
Copied to clipboard
| Challenge: | Multilingual word embeddings embed words from many languages into a single semantic space such that words with similar meanings are close to each other regardless of the language. |
| Approach: | They propose to use multilingual word embeddings to align embeddable words from multiple languages into a single semantic space so that words with similar meanings are close to each other regardless of the language. |
| Outcome: | The proposed model can be used to learn gender bias in multilingual representations and to improve transfer learning. |
Stereotypes and Smut: The (Mis)representation of Non-cisgender Identities by Text-to-Image Models (2023.findings-acl)
Copied to clipboard
| Challenge: | Initial studies have pointed to the potential for harm due to predictive bias, reflecting and potentially reinforcing cultural stereotypes. |
| Approach: | They conduct a survey among non-cisgender individuals and interviews to establish which harms affected individuals anticipate, and how they would like to be represented. |
| Outcome: | The results show that certain non-cisgender identities are consistently (mis)represented as less human, more stereotyped and more sexualised. |
When Cultures Meet: Multicultural Text-to-Image Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | a new task to evaluate text-to-image generation models for multicultural scenes is unexplored. |
| Approach: | They propose a benchmark task to evaluate text-to-image models in multicultural settings . they use a dataset of 9,000 images spanning five countries, three age groups, two genders, 25 historical landmarks, and five languages to analyze behavior . |
| Outcome: | The proposed benchmark analyzes the behavior of state-of-the-art models across multiple dimensions including alignment, image quality, aesthetics, knowledge, and fairness. |
How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions? (2022.emnlp-main)
Copied to clipboard
| Challenge: | Text-to-image generative models can generate high-quality photo-realistic images conditional on natural language text descriptions in a zero-shot fashion. |
| Approach: | They propose an Ethical NaTural Language Interventions in Text-to-Image GENeration benchmark dataset to evaluate the change in image generation conditional on ethical interventions across three social axes – gender, skin color, and culture. |
| Outcome: | The proposed model generations cover diverse social groups while preserving image quality. |
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation (2024.acl-long)
Copied to clipboard
Akshita Jha, Vinodkumar Prabhakaran, Remi Denton, Sarah Laszlo, Shachi Dave, Rida Qadri, Chandan Reddy, Sunipa Dev
| Challenge: | Existing approaches for evaluating stereotypes have a noticeable lack of coverage of global identity groups and their associated stereotypes. |
| Approach: | They propose to use a dataset to evaluate nationality-based stereotypes in T2I models across 135 nationalities to assess offensive stereotypes. |
| Outcome: | The proposed dataset enables evaluation of known nationality-based stereotypes across 135 nationalities. |
Social Bias in Multilingual Language Models: A Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | Pretrained multilingual models exhibit the same social bias as models processing English texts. |
| Approach: | They examine the literature on bias evaluation and mitigation approaches in multilingual and non-English contexts and identify gaps in the field. |
| Outcome: | The proposed models perform well on multilingual language understanding benchmarks and are consistent with the current literature. |
On Evaluating and Mitigating Gender Biases in Multilingual Settings (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks and resources for evaluating gender biases in multilingual settings are limited. |
| Approach: | They propose to extend DisCo to different Indian languages using human annotations to evaluate gender biases in multilingual models. |
| Outcome: | The proposed benchmarks and mitigation techniques are extended beyond English to evaluate gender biases in multilingual models. |