Challenge: Large-scale Vision-Language Models (LVLMs) are being deployed in real-world settings that require visual inference.
Approach: They evaluate LVLMs' ability to account for variation in color perception using the Ishihara Test.
Outcome: The proposed models fail to reproduce the perceptual outcomes experienced by affected individuals and default to normative color perception.

Similar Papers

Rainbow - A Benchmark for Systematic Testing of How Sensitive Visio-Linguistic Models are to Color Naming (2024.eacl-long)

Copied to clipboard

Challenge: Visio-linguistic models have been gaining popularity for tasks that require a deeper understanding of multimodalities.
Approach: They compile a probing dataset to test multi-modal alignment around color . they show that models have trouble with prepositions and verbs .
Outcome: The proposed model is superior to models that do not rely on pre-extracted image features and is able to perform well with noisy pre-training data.
Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward (2025.findings-acl)

Copied to clipboard

Challenge: Large Vision Language Models (LVLMs) have shown impressive performance on various vision-language tasks.
Approach: They propose a benchmark framework for evaluating Visual Variation Robustness of Large Vision Language Models that incorporates automated evaluation dataset generation and principled metrics for thorough robustness assessment.
Outcome: The proposed framework identifies a vulnerability to visual variations affecting even advanced models that excel at complex vision-language tasks but significantly underperform on simple tasks like object recognition.
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in instruction-tuned Large Vision-Language Models (LVLMs) have imbued the models with the ability to generate high-level, image-grounded explanations with ease.
Approach: They propose to use a multiple granularity attribute-centric benchmark and training mixture to evaluate LVLMs’ fine-grained visual comprehension ability.
Outcome: The proposed model improves on LLaVa-1.5, InstructBLIP and GPT-4V and demonstrates that they struggle to generate descriptive visual attributes based on a concept that appears within an input image despite their prominent zero-shot image captioning ability.
Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model (2025.acl-long)

Copied to clipboard

Challenge: Existing models for large vision-language tasks are trained on English data, which makes them struggle to understand non-English input and fail to generate output in the desired target language.
Approach: They conduct multi-stage experiments on 13 vision-language tasks and 43 languages . they find that one can include as many as 100 training languages simultaneously with as little as 25-50% of non-English data .
Outcome: The proposed model outperforms existing models in 14 tasks and 56 languages.
MVP-Bench: Can Large Vision-Language Models Conduct Multi-level Visual Perception Like Humans? (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing LVLMs perform visual perception at multiple levels, but they are not able to perform multi-level tasks.
Approach: They propose a visual–language benchmark to evaluate LVLMs' perceptions . they use manipulated images to examine how LVLs can perform multi-level tasks .
Outcome: The proposed model performs poorly on high-level perception tasks, the authors show . they also show that current models do not generalize in understanding semantics of synthetic images .
Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals (2025.naacl-long)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) have been proposed to augment LLMs with visual inputs.
Approach: They propose large vision-Language Models to augment LLMs with visual inputs.
Outcome: The proposed models condition generated text on both an input image and a visual prompt, enabling a variety of use cases such as visual question answering and multimodal chat.
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models (2024.eacl-long)

Copied to clipboard

Challenge: Existing vision-and-language models perform better on multimodal tasks, but there is little understanding of how multimodal learning can help visual representations.
Approach: They conduct a probing analysis of visual representations in existing vision-and-language models and vision-only models by probing on a broad range of tasks.
Outcome: The proposed model improves vision-and-language models on label and attribute prediction tasks while vision-only models are stronger on dense prediction tasks.
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities (2025.acl-long)

Copied to clipboard

Challenge: Vision-language Models have been shown to be highly capable but lacking basic visual understanding skills.
Approach: They propose to examine the limitations of vision-language models on visual tasks by constructing a series of tests that probe which components of design may be lacking.
Outcome: The proposed tests compare VLMs to other models on visual encoders, intermediate vision-language projection and LLM-decoder outputs.
The World of an Octopus: How Reporting Bias Influences a Language Model’s Perception of Color (2021.emnlp-main)

Copied to clipboard

Challenge: Recent work has raised concerns about the inherent limitations of text-only pretraining.
Approach: They first generate a color dataset of human-perceived color distributions for 521 common objects and then use it to analyze and compare the color distribution found in text and the distribution captured by language models.
Outcome: The proposed model improves on the CoDa color distribution, while the language model improve on the ground-truth distribution.
Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models (2026.findings-acl)

Copied to clipboard

Challenge: a recent study shows that vision-language models have modality gaps that persist even in well-aligned models.
Approach: They propose a modality-dominance score to measure and leverage modality gaps . they propose automatic interpretability metrics to evaluate these features in a scalable manner .
Outcome: The proposed framework allows for training-free probing and editing methods for understanding model perception across genders and generating adversarial examples.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations