How Well Do Vision Models Encode Diagram Attributes? (2024.acl-srw)

Copied to clipboard

Challenge: Experimental results show vision models struggle to identify diagram attributes such as node colors and shapes, along with edge colors and connection patterns.
Approach: They evaluated vision models and retrieving diagrams using text queries to determine how well they recognize diagram attributes and edge connection patterns.
Outcome: The models can recognize node colors, shapes, and edge colors, but struggle to identify differences in edge connection patterns that play a pivotal role in the semantics of diagrams.

Similar Papers

Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder (2025.acl-long)

Copied to clipboard

Challenge: Recent studies show that CLIP models struggle with visual reasoning tasks . despite the success of Contrastive Language-Image Pretraining, there are still limitations .
Approach: They propose to use a visual encoder to train CLIP-like models for fine-grained visual reasoning tasks.
Outcome: The proposed models outperform CLIP-like encoders in visual reasoning tasks . the study highlights the importance of VLM architectural choices .
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that deep vision-only and language-only models project inputs into a partially aligned representational space.
Approach: They investigate whether a model's representational code is semantically shared . they find that alignment peaks in mid-to-late layers of both model types .
Outcome: a forced-choice "Pick-a-Pic" task shows human preferences for image-caption matches are mirrored in embedding spaces across vision-language model pairs.
Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions (2025.acl-long)

Copied to clipboard

Challenge: Existing studies show that direct generation of diagram descriptions is costly and biased against blind and low-vision (BLV) users.
Approach: They ask sighted individuals to assess diagram descriptions generated by vision-language models . they use latent supervision to guide the models with latent inference .
Outcome: The results show that visual descriptions generated by vision-language models are effective and useful to educators who are themselves BLV and teach visually impaired learners.
Do Vision-Language Models Understand Compound Nouns? (2024.naacl-short)

Copied to clipboard

Challenge: Open-vocabulary vision-language models (CLIP) are emerging as a promising new paradigm for text-to-image retrieval.
Approach: They propose a benchmark to evaluate the effectiveness of open-vocabulary vision-language models (CLIP) for text-to-image retrieval using contrastive loss.
Outcome: The proposed framework improves CN understanding of CLIP by 8.25% on Compun.
If CLIP Could Talk: Understanding Vision-Language Model Representations Through Their Preferred Concept Descriptions (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies assume that VLMs prioritize visual attributes to represent concepts.
Approach: They propose a novel approach to characterize features important for VLMs using reinforcement learning.
Outcome: The proposed approach characterizes features that are important for VLMs . it shows that spurious descriptions have a major role in VLM representations despite providing no helpful information.
Probing Contextual Language Models for Common Ground with Visual Representations (2021.naacl-main)

Copied to clipboard

Challenge: Contextual language models have attracted great interest in probing what is encoded in their representations.
Approach: They propose a probing model that evaluates how effective are text-only representations in distinguishing between matching and non-matching visual representations.
Outcome: The proposed model outperforms text-only language models in instance retrieval, but underperform humans.
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities (2025.acl-long)

Copied to clipboard

Challenge: Vision-language Models have been shown to be highly capable but lacking basic visual understanding skills.
Approach: They propose to examine the limitations of vision-language models on visual tasks by constructing a series of tests that probe which components of design may be lacking.
Outcome: The proposed tests compare VLMs to other models on visual encoders, intermediate vision-language projection and LLM-decoder outputs.
Can VLMs Actually See and Read? A Survey on Modality Collapse in Vision-Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Vision-language models integrate textual and visual information, enabling them to process visual inputs and generate predictions.
Approach: They review work on modality collapse analysis to provide insights into the reason for this unintended behavior and review probing studies for fine-grained vision-language understanding.
Outcome: The proposed models can achieve competitive performance in vision-language tasks despite relying heavily on textual information and ignoring visual information.
Delving into the Openness of CLIP (2023.findings-acl)

Copied to clipboard

Challenge: Contrastive Language-Image Pre-training (CLIP) allows for open-vocabulary visual recognition, where the model can recognize images from an open class set in a zero-shot manner.
Approach: They propose to use image classification as an image-to-text matching task instead of discrete category IDs to achieve open-vocabulary visual recognition.
Outcome: The proposed model can recognize images from an open vocabulary in a zero-shot manner, but its performance deteriorates as the vocabulary expands.
Paparazzi: A Deep Dive into the Capabilities of Language and Vision Models for Grounding Viewpoint Descriptions (2023.findings-eacl)

Copied to clipboard

Challenge: Existing language and vision models can be used for language understanding in 3D environments . however, existing models lack specific properties and biases that limit their performance .
Approach: They propose a framework that uses a camera to generate images from different viewpoints and evaluate them in terms of their similarity to natural language descriptions.
Outcome: The proposed model performs poorly on most canonical views and fine-tunes using hard negative sampling and random contrasting yields good results even under conditions with little available training data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations