Challenge: Recent work has raised concerns about the inherent limitations of text-only pretraining.
Approach: They first generate a color dataset of human-perceived color distributions for 521 common objects and then use it to analyze and compare the color distribution found in text and the distribution captured by language models.
Outcome: The proposed model improves on the CoDa color distribution, while the language model improve on the ground-truth distribution.

Similar Papers

Do ever larger octopi still amplify reporting biases? Evidence from judgments of typical colour (2022.aacl-short)

Copied to clipboard

Challenge: Language models trained on text-only corpora have no direct access to the physical world and thus suffer from reporting bias.
Approach: They investigate reporting bias from the perspective of colour in larger language models such as PaLM and GPT-3.
Outcome: The proposed models outperform smaller models on the basis of colour and more closely track human judgements than smaller models.
Multi-Modal Bias: Introducing a Framework for Stereotypical Bias Assessment beyond Gender and Race in Vision–Language Models (2023.eacl-main)

Copied to clipboard

Challenge: Recent advances in self-supervised training have led to a new class of pretrained vision–language models.
Approach: They propose a visual and textual bias benchmark to assess bias in self-supervised multimodal models using 3,800 images and phrases from 14 population subgroups.
Outcome: The proposed model shows that it favors certain groups while maintaining the accuracy of the model.
Do Neural Language Models Overcome Reporting Bias? (2020.coling-main)

Copied to clipboard

Challenge: Recent studies show that pre-trained language models can overcome reporting bias by estimating the plausibility of rare but unspoken facts.
Approach: They revisit the experiments conducted by Gordon and Van Durme (2013) . they find that pre-trained language models overestimate the very rare .
Outcome: The proposed approach overestimates the rare at the expense of the rare, while minimizing reporting bias.
Visual Commonsense in Pretrained Unimodal and Multimodal Models (2022.naacl-main)

Copied to clipboard

Challenge: Fig. 1 shows how text-only and image-only models can capture commonsense visual attributes, but reporting bias affects their performance.
Approach: They use a Visual Commonsense Tests dataset to validate their findings . they find multimodal models better reconstruct attribute distributions, but are still subject to reporting bias .
Outcome: The proposed model improves on the unimodal and multimodal models, but is still subject to reporting bias.
Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders (2025.naacl-long)

Copied to clipboard

Challenge: Recent work has found that vision-language models trained under the Contrastive Language Image Pre-training framework contain intrinsic social biases, but how these biase relates to downstream performance has been unclear.
Approach: They present the largest comprehensive analysis to-date of how upstream pre-training factors and downstream performance of CLIP models relate to their intrinsic biases.
Outcome: The proposed model performance analysis shows that the choice of pre-training dataset is the most significant upstream predictor of bias, whereas architectural variations have minimal impact.
What do Models Learn From Training on More Than Text? Measuring Visual Commonsense Knowledge (2022.acl-srw)

Copied to clipboard

Challenge: Existing evaluation methods to measure what language models learn from multimodal training are lacking.
Approach: They propose two evaluation tasks to measure commonsense knowledge in language models by using visual data to evaluate multimodal models and unimodal baselines.
Outcome: The proposed evaluation tasks show that training on a visual modality improves on the visual commonsense knowledge in language models.
A Prompt Array Keeps the Bias Away: Debiasing Vision-Language Models with Adversarial Learning (2022.aacl-main)

Copied to clipboard

Challenge: Large-scale, pretrained vision-language models are growing in popularity due to impressive performance on downstream tasks with minimal finetuning.
Approach: They propose to apply ranking metrics to image-text representations to investigate bias measures and debiasing methods to reduce various bias measures.
Outcome: The proposed model reduces bias measures with minimal degradation to image-text representations.
Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals (2025.naacl-long)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) have been proposed to augment LLMs with visual inputs.
Approach: They propose large vision-Language Models to augment LLMs with visual inputs.
Outcome: The proposed models condition generated text on both an input image and a visual prompt, enabling a variety of use cases such as visual question answering and multimodal chat.
Assessing Multilingual Fairness in Pre-trained Multimodal Representations (2022.findings-acl)

Copied to clipboard

Challenge: Recent pre-trained multimodal models have shown exceptional capabilities towards connecting images and natural language.
Approach: They propose two new fairness notions for pre-trained multimodal models that consider language as the fairness recipient.
Outcome: The proposed models can be generalized to multilingualism by cross-lingual alignment . the results show that the models are individually fair across languages .
Bias at a Second Glance: A Deep Dive into Bias for German Educational Peer-Review Data Modeling (2022.coling-1)

Copied to clipboard

Challenge: Existing studies have highlighted a variety of biases in pre-trained language models . however, these studies focus on fine-grained analysis of educational corpora and text that is not English .
Approach: They analyze bias across text and through multiple architectures on a corpus of 9,165 German peer-reviews collected from university students over five years.
Outcome: The proposed dataset shows that pre-trained language models exhibit conceptual, racial, and gender biases.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations