Challenge: Despite significant progress in multimodal language models, it remains unclear whether visual grounding enhances their understanding of embodied knowledge compared to text-only models.
Approach: They propose to assess vision-language models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions.
Outcome: The proposed benchmark assesses the models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions.

Similar Papers

Can Language Models Understand Physical Concepts? (2023.emnlp-main)

Copied to clipboard

Challenge: Existing language models do not understand basic physical concepts in the human world.
Approach: They propose a method to transfer embodied knowledge from visual models to LMs . they use visual concepts and embodies concepts learned from interaction with the world .
Outcome: The proposed method achieves comparable performance with scaling up parameters of LMs 134.
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces (2025.acl-long)

Copied to clipboard

Challenge: Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban aerial spaces remain to be explored.
Approach: They propose a benchmark to evaluate whether large multimodal models can process continuous first-person visual observations like humans.
Outcome: The proposed model can process first-person visual observations like humans, enabling recall, perception, reasoning, and navigation.
Visual Grounding Helps Learn Word Meanings in Low-Data Regimes (2024.naacl-long)

Copied to clipboard

Challenge: Modern neural language models (LMs) require distinctly un-human-like ways to achieve these results.
Approach: They train a diverse set of LM architectures with and without auxiliary visual supervision on datasets of varying scales.
Outcome: The proposed models exhibit better learning of syntactic categories, lexical relations, semantic features, word similarity and alignment with human neural representations.
Multimodal Language Models Show Evidence of Embodied Simulation (2024.lrec-main)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) are gaining popularity as partial solutions to the “symbol grounding problem” faced by language models trained on text alone.
Approach: They propose to use multimodal large language models to integrate linguistic representations with data from other modalities to investigate whether they are integrated into a model.
Outcome: The proposed models are sensitive to visual features like object shape when it is implied by a verbal description of an event.
Does Vision-and-Language Pretraining Improve Lexical Grounding? (2021.findings-emnlp)

Copied to clipboard

Challenge: Large pretrained language models (LMs) have been criticized for lack of grounding, i.e., connecting words to their meanings in the physical world.
Approach: They compare vision-and-language (VL) models trained jointly on text and image or video data to find out how they compare to text-only counterparts.
Outcome: The proposed model outperforms the text-only variants on a commonsense question answering task.
Does Vision Accelerate Hierarchical Generalization in Neural Language Learners? (2025.coling-main)

Copied to clipboard

Challenge: Neural language models (LMs) are arguably less data-efficient than humans from a language acquisition perspective.
Approach: They investigate the advantage of grounded language acquisition over visual input to improve syntactic generalization.
Outcome: The proposed model is less efficient than humans in language acquisition . it shows that visual input helps syntactic generalization, but not vision .
Exploring Spatial Schema Intuitions in Large Language and Vision Models (2024.findings-acl)

Copied to clipboard

Challenge: Large language models excel in varied NLP tasks, but lack a direct connection between sensory perception and physical action.
Approach: They examine whether large language models capture implicit human intuitions about building blocks of language . they employ spatial cognitive foundations developed through early sensorimotor experiences .
Outcome: The proposed model captures implicit human intuitions about building blocks of language without a tangible connection to embodied experiences.
What do Models Learn From Training on More Than Text? Measuring Visual Commonsense Knowledge (2022.acl-srw)

Copied to clipboard

Challenge: Existing evaluation methods to measure what language models learn from multimodal training are lacking.
Approach: They propose two evaluation tasks to measure commonsense knowledge in language models by using visual data to evaluate multimodal models and unimodal baselines.
Outcome: The proposed evaluation tasks show that training on a visual modality improves on the visual commonsense knowledge in language models.
Grounding Meaning Representation for Situated Reasoning (2022.aacl-tutorials)

Copied to clipboard

Challenge: a tutorial aims to build agents that understand language using a simulated environment . situated reasoning is a critical aspect of human language understanding .
Approach: This tutorial combines a synthesis of multimodal grounding and meaning representation techniques with formal and computational models of situated reasoning.
Outcome: This tutorial combines multimodal grounding and meaning representation techniques with formal and computational models of embodied reasoning.
LMs stand their Ground: Investigating the Effect of Embodiment in Figurative Language Interpretation by Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Figures are based on the use of words in a way that deviates from their conventional order and meaning.
Approach: They propose to use a figurative language model to interpret embodied metaphors by using larger language models that conceptualise embodies the action of the metaphorical sentence.
Outcome: The proposed model enables interpretation of figurative language when the action of the metaphorical sentence is more embodied.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations