Challenge: Recent advances in self-supervised training have led to a new class of pretrained vision–language models.
Approach: They propose a visual and textual bias benchmark to assess bias in self-supervised multimodal models using 3,800 images and phrases from 14 population subgroups.
Outcome: The proposed model shows that it favors certain groups while maintaining the accuracy of the model.

Similar Papers

A Multi-dimensional study on Bias in Vision-Language models (2023.findings-acl)

Copied to clipboard

Challenge: Recent studies have focused on the issue of bias in joint Vision-Language models . pre-trained models complete a neutral template with a hurtful word 5% of the time .
Approach: They propose to use a multi-dimensional bias metric to investigate bias in English VL models . they use gender, ethnicity, and age as dimensions to analyze bias in VLs .
Outcome: The proposed model is based on gender, ethnicity, and age as dimensions.
ModSCAN: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language Modalities (2024.emnlp-main)

Copied to clipboard

Challenge: Large vision-language models have been widely used but stereotypical biases are unexplored.
Approach: They propose a framework to SCAN stereotypical bias within large vision-language models . they examine stereotype biases with respect to gender and race in three scenarios .
Outcome: The proposed framework can reduce stereotypical biases in large vision-language models . the currently popular models show significant stereotype biase .
Social Bias in Multilingual Language Models: A Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Pretrained multilingual models exhibit the same social bias as models processing English texts.
Approach: They examine the literature on bias evaluation and mitigation approaches in multilingual and non-English contexts and identify gaps in the field.
Outcome: The proposed models perform well on multilingual language understanding benchmarks and are consistent with the current literature.
VLStereoSet: A Study of Stereotypical Bias in Pre-trained Vision-Language Models (2022.aacl-main)

Copied to clipboard

Challenge: Existing studies on pre-trained vision-language models have focused on measuring biases and stereotypes in a single modality.
Approach: They extend a recently released stereotypical bias dataset into a vision-language probing dataset called VLStereoSet to measure stereotypical biased vision-linguistic models.
Outcome: The proposed probing task measures stereotypical bias in vision-language models and its intra-modal and inter-modal biases.
A Prompt Array Keeps the Bias Away: Debiasing Vision-Language Models with Adversarial Learning (2022.aacl-main)

Copied to clipboard

Challenge: Large-scale, pretrained vision-language models are growing in popularity due to impressive performance on downstream tasks with minimal finetuning.
Approach: They propose to apply ranking metrics to image-text representations to investigate bias measures and debiasing methods to reduce various bias measures.
Outcome: The proposed model reduces bias measures with minimal degradation to image-text representations.
Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals (2025.naacl-long)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) have been proposed to augment LLMs with visual inputs.
Approach: They propose large vision-Language Models to augment LLMs with visual inputs.
Outcome: The proposed models condition generated text on both an input image and a visual prompt, enabling a variety of use cases such as visual question answering and multimodal chat.
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have highlighted the existence of social biases within large vision and language models.
Approach: They propose a framework for systematically evaluating gender, race, and age biases in vision-language models with respect to professions.
Outcome: The proposed framework covers all supported inference modes of the recent vision-language models, including image-to-text, text-to image, and image- to-image.
More than Minorities and Majorities: Understanding Multilateral Bias in Language Generation (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies on bias dataset construction and mitigation focus on one demographic group . in real-world applications, there are more than two demographic groups at risk of the same bias.
Approach: They propose to analyze and reduce biases across multiple demographic groups using a multi-demographic bias dataset.
Outcome: The proposed method can mitigate biases among multiple demographic groups effectively, the authors show .
Seeing Race, Feeling Bias: Emotion Stereotyping in Multimodal Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Emotion stereotypes are also tightly tied to race and skin tone, but previous studies have overlooked this dimension.
Approach: They propose a multimodal study of racial, gender, and skin-tone bias in emotion attribution . they evaluate four open-source MLLMs using 2.1K emotion-related events .
Outcome: The proposed study examines four open-source MLLMs using 2.1K emotion-related events paired with 400 neutral face images across three different prompt strategies.
Addressing Healthcare-related Racial and LGBTQ+ Biases in Pretrained Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Pretrained language models (PLMs) propagate social stigmas and stereotypes, a critical concern given their widespread use.
Approach: They adapt two intrinsic bias benchmarks to quantify racial and LGBTQ+ biases in prevalent PLMs and empirically evaluate the effectiveness of various debiasing methods in mitigating these biase.
Outcome: The proposed methods reduce biases without compromising performance in downstream tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations