Challenge: Existing language and vision models can be used for language understanding in 3D environments . however, existing models lack specific properties and biases that limit their performance .
Approach: They propose a framework that uses a camera to generate images from different viewpoints and evaluate them in terms of their similarity to natural language descriptions.
Outcome: The proposed model performs poorly on most canonical views and fine-tunes using hard negative sampling and random contrasting yields good results even under conditions with little available training data.

Similar Papers

What Does BERT with Vision Look At? (2020.acl-main)

Copied to clipboard

Challenge: Pre-trained visual grounded language models have improved performance on vision-and-language tasks but what they learn during pre-training remains unclear.
Approach: They show that attention heads of visual grounded language models actively ground elements of language to image regions.
Outcome: The attention heads of a visual grounded language model can ground elements to image regions, demonstrating their ability to detect syntactic relations between non-entity words and image regions.
Probing Contextual Language Models for Common Ground with Visual Representations (2021.naacl-main)

Copied to clipboard

Challenge: Contextual language models have attracted great interest in probing what is encoded in their representations.
Approach: They propose a probing model that evaluates how effective are text-only representations in distinguishing between matching and non-matching visual representations.
Outcome: The proposed model outperforms text-only language models in instance retrieval, but underperform humans.
RWKV-CLIP: A Robust Vision-Language Representation Learner (2024.emnlp-main)

Copied to clipboard

Challenge: Using large image-text datasets, large-scale image-data sets have been used for visionlanguage pre-training.
Approach: They propose a framework that leverages Large Language Models to combine and refine information from web-based image-text pairs, synthetic captions, and detection tags.
Outcome: The proposed framework can combine and refine information from web-based image-text pairs, synthetic captions, and detection tags.
Visually Grounded Reasoning across Languages and Cultures (2021.emnlp-main)

Copied to clipboard

Challenge: a new protocol allows for a multilingual hierarchy of concepts and images based on native speakers . the results suggest that the current models are not robust enough to handle multilingual data .
Approach: They propose a protocol to construct an ImageNet-style hierarchy representative of more languages and cultures.
Outcome: The proposed protocol lets the selection of concepts and images be entirely driven by native speakers, rather than scraping them automatically.
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder (2025.acl-long)

Copied to clipboard

Challenge: Recent studies show that CLIP models struggle with visual reasoning tasks . despite the success of Contrastive Language-Image Pretraining, there are still limitations .
Approach: They propose to use a visual encoder to train CLIP-like models for fine-grained visual reasoning tasks.
Outcome: The proposed models outperform CLIP-like encoders in visual reasoning tasks . the study highlights the importance of VLM architectural choices .
Read Before Grounding: Scene Knowledge Visual Grounding via Multi-step Parsing (2025.coling-main)

Copied to clipboard

Challenge: Existing VG datasets use simple textual descriptions with limited attribute and spatial information between images and text.
Approach: They propose a method that transforms visual knowledge into concise, information-dense visual descriptions.
Outcome: The proposed method significantly improves performance of multimodal grounding models.
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that deep vision-only and language-only models project inputs into a partially aligned representational space.
Approach: They investigate whether a model's representational code is semantically shared . they find that alignment peaks in mid-to-late layers of both model types .
Outcome: a forced-choice "Pick-a-Pic" task shows human preferences for image-caption matches are mirrored in embedding spaces across vision-language model pairs.
Multimodal Pretraining Unmasked: A Meta-Analysis and a Unified Framework of Vision-and-Language BERTs (2021.tacl-1)

Copied to clipboard

Challenge: Large-scale pretraining and task-specific fine-tuning are now the standard methodology for many tasks in computer vision and natural language processing.
Approach: They propose to combine two types of vision and language BERTs to create a theoretical framework that can be unified under different theoretical frameworks.
Outcome: The proposed models can be classified into single-stream or dual-stream encoders and are unified under a single theoretical framework.
Language in a (Search) Box: Grounding Language Learning in Real-World Human-Machine Interaction (2021.naacl-main)

Copied to clipboard

Challenge: Scholarly work in this area uses toy worlds and synthetic linguistic data, but grounded language learning offers several practical and scientific advantages.
Approach: They propose to model teacher-learner dynamics through natural interactions occurring between users and search engines.
Outcome: The proposed model is better than non-grounded models on compositionality and zero-shot inference tasks.
Delving into the Openness of CLIP (2023.findings-acl)

Copied to clipboard

Challenge: Contrastive Language-Image Pre-training (CLIP) allows for open-vocabulary visual recognition, where the model can recognize images from an open class set in a zero-shot manner.
Approach: They propose to use image classification as an image-to-text matching task instead of discrete category IDs to achieve open-vocabulary visual recognition.
Outcome: The proposed model can recognize images from an open vocabulary in a zero-shot manner, but its performance deteriorates as the vocabulary expands.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations