Challenge: Existing studies have highlighted the existence of social biases within large vision and language models.
Approach: They propose a framework for systematically evaluating gender, race, and age biases in vision-language models with respect to professions.
Outcome: The proposed framework covers all supported inference modes of the recent vision-language models, including image-to-text, text-to image, and image- to-image.

Similar Papers

VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing studies on VLM bias focus on portrait-style images and gender-occupation associations . existing studies ignore broader and more complex social stereotypes and their implied harm .
Approach: They propose a large-scale VQA benchmark for evaluating bias in vision-language models . they use a question-answering framework that spans factuality, perception, stereotyping, and decision making .
Outcome: The proposed framework examines bias in vision-language models using 30M+ images . findings reveal subtle, multifaceted, and surprising stereotypical patterns .
BiasDora: Exploring Hidden Biased Associations in Vision-Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on social biases focus on a limited set of documented associations, such as gender-profession or race-crime.
Approach: They propose to examine hidden, implicit bias associations across 9 bias dimensions by probing VLMs to uncover hidden, unexamined associations.
Outcome: The proposed methods reveal that biases vary in negativity, toxicity, and extremity.
Multi-Modal Bias: Introducing a Framework for Stereotypical Bias Assessment beyond Gender and Race in Vision–Language Models (2023.eacl-main)

Copied to clipboard

Challenge: Recent advances in self-supervised training have led to a new class of pretrained vision–language models.
Approach: They propose a visual and textual bias benchmark to assess bias in self-supervised multimodal models using 3,800 images and phrases from 14 population subgroups.
Outcome: The proposed model shows that it favors certain groups while maintaining the accuracy of the model.
Examining Gender and Racial Bias in Large Vision–Language Models Using a Novel Dataset of Parallel Images (2024.eacl-long)

Copied to clipboard

Challenge: a new wave of large vision–language models (LVLMs) incorporate images as input in addition to text . a recent study examined potential gender and racial biases in such systems based on the perceived characteristics of the people in the input images.
Approach: They examine potential gender and racial biases in large vision–language models . they query a dataset of AI-generated images of people to see whether they differ .
Outcome: The proposed dataset shows that the images differ in gender and race according to the perceived characteristics of the person depicted.
Ask Me Again Differently: GRAS for Measuring Bias in Vision Language Models on Gender, Race, Age, and Skin Tone (2026.findings-eacl)

Copied to clipboard

Challenge: Using vision language models, we examine demographic biases in VLMs across gender, race, age, and skin tone.
Approach: They propose a benchmark for uncovering demographic biases in Vision Language Models . they propose 'Gras Bias Score' to quantify bias in VLMs based on gender, race, age and skin tone .
Outcome: The proposed model achieves 98, far from the unbiased ideal of 0.
ModSCAN: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language Modalities (2024.emnlp-main)

Copied to clipboard

Challenge: Large vision-language models have been widely used but stereotypical biases are unexplored.
Approach: They propose a framework to SCAN stereotypical bias within large vision-language models . they examine stereotype biases with respect to gender and race in three scenarios .
Outcome: The proposed framework can reduce stereotypical biases in large vision-language models . the currently popular models show significant stereotype biase .
Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos (2026.acl-long)

Copied to clipboard

Challenge: Vision-Language Models (VLMs) are increasingly deployed in socially consequential settings . attribution under visual confounding is a central challenge in measuring social bias .
Approach: They propose a face-only counterfactual evaluation paradigm that isolates demographic effects while preserving real-image realism.
Outcome: The proposed paradigm isolates demographic effects while preserving real-image realism.
Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals (2025.naacl-long)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) have been proposed to augment LLMs with visual inputs.
Approach: They propose large vision-Language Models to augment LLMs with visual inputs.
Outcome: The proposed models condition generated text on both an input image and a visual prompt, enabling a variety of use cases such as visual question answering and multimodal chat.
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts (2025.findings-emnlp)

Copied to clipboard

Challenge: Large vision-language models have demonstrated strong capabilities in open-world visual understanding, but it is not clear how they address demographic biases in real life.
Approach: They propose a method to assess visual fairness in LVLMs by question-answering/classification tasks.
Outcome: The proposed approach improves transparency and offers a scalable solution for fairness mitigation.
Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes (2025.findings-acl)

Copied to clipboard

Challenge: Existing encoder-based vision-language models (VLMs) contain intrinsic biases that manifest in biased outputs.
Approach: They propose a framework to measure intrinsic bias propagation by correlating intrinsic bias with extrinsic bias in zero-shot text-to-image and image-totext retrieval.
Outcome: The proposed framework shows that larger/better-performing models exhibit greater bias propagation, raising concerns given the trend towards increasingly complex AI models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations