Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Despite significant progress in multimodal language models, it remains unclear whether visual grounding enhances their understanding of embodied knowledge compared to text-only models. |
| Approach: | They propose to assess vision-language models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions. |
| Outcome: | The proposed benchmark assesses the models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions. |
Similar Papers
Can Language Models Understand Physical Concepts? (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing language models do not understand basic physical concepts in the human world. |
| Approach: | They propose a method to transfer embodied knowledge from visual models to LMs . they use visual concepts and embodies concepts learned from interaction with the world . |
| Outcome: | The proposed method achieves comparable performance with scaling up parameters of LMs 134. |
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces (2025.acl-long)
Copied to clipboard
Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang, Jirong Zha, Weichen Zhang, Chen Gao, Yue Wang, Jinqiang Cui, Xinlei Chen, Yong Li
| Challenge: | Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban aerial spaces remain to be explored. |
| Approach: | They propose a benchmark to evaluate whether large multimodal models can process continuous first-person visual observations like humans. |
| Outcome: | The proposed model can process first-person visual observations like humans, enabling recall, perception, reasoning, and navigation. |
Visual Grounding Helps Learn Word Meanings in Low-Data Regimes (2024.naacl-long)
Copied to clipboard
| Challenge: | Modern neural language models (LMs) require distinctly un-human-like ways to achieve these results. |
| Approach: | They train a diverse set of LM architectures with and without auxiliary visual supervision on datasets of varying scales. |
| Outcome: | The proposed models exhibit better learning of syntactic categories, lexical relations, semantic features, word similarity and alignment with human neural representations. |
Multimodal Language Models Show Evidence of Embodied Simulation (2024.lrec-main)
Copied to clipboard
| Challenge: | Multimodal large language models (MLLMs) are gaining popularity as partial solutions to the “symbol grounding problem” faced by language models trained on text alone. |
| Approach: | They propose to use multimodal large language models to integrate linguistic representations with data from other modalities to investigate whether they are integrated into a model. |
| Outcome: | The proposed models are sensitive to visual features like object shape when it is implied by a verbal description of an event. |
Does Vision-and-Language Pretraining Improve Lexical Grounding? (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Large pretrained language models (LMs) have been criticized for lack of grounding, i.e., connecting words to their meanings in the physical world. |
| Approach: | They compare vision-and-language (VL) models trained jointly on text and image or video data to find out how they compare to text-only counterparts. |
| Outcome: | The proposed model outperforms the text-only variants on a commonsense question answering task. |
Does Vision Accelerate Hierarchical Generalization in Neural Language Learners? (2025.coling-main)
Copied to clipboard
| Challenge: | Neural language models (LMs) are arguably less data-efficient than humans from a language acquisition perspective. |
| Approach: | They investigate the advantage of grounded language acquisition over visual input to improve syntactic generalization. |
| Outcome: | The proposed model is less efficient than humans in language acquisition . it shows that visual input helps syntactic generalization, but not vision . |
Exploring Spatial Schema Intuitions in Large Language and Vision Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Large language models excel in varied NLP tasks, but lack a direct connection between sensory perception and physical action. |
| Approach: | They examine whether large language models capture implicit human intuitions about building blocks of language . they employ spatial cognitive foundations developed through early sensorimotor experiences . |
| Outcome: | The proposed model captures implicit human intuitions about building blocks of language without a tangible connection to embodied experiences. |
What do Models Learn From Training on More Than Text? Measuring Visual Commonsense Knowledge (2022.acl-srw)
Copied to clipboard
| Challenge: | Existing evaluation methods to measure what language models learn from multimodal training are lacking. |
| Approach: | They propose two evaluation tasks to measure commonsense knowledge in language models by using visual data to evaluate multimodal models and unimodal baselines. |
| Outcome: | The proposed evaluation tasks show that training on a visual modality improves on the visual commonsense knowledge in language models. |
Grounding Meaning Representation for Situated Reasoning (2022.aacl-tutorials)
Copied to clipboard
| Challenge: | a tutorial aims to build agents that understand language using a simulated environment . situated reasoning is a critical aspect of human language understanding . |
| Approach: | This tutorial combines a synthesis of multimodal grounding and meaning representation techniques with formal and computational models of situated reasoning. |
| Outcome: | This tutorial combines multimodal grounding and meaning representation techniques with formal and computational models of embodied reasoning. |
LMs stand their Ground: Investigating the Effect of Embodiment in Figurative Language Interpretation by Language Models (2023.findings-acl)
Copied to clipboard
| Challenge: | Figures are based on the use of words in a way that deviates from their conventional order and meaning. |
| Approach: | They propose to use a figurative language model to interpret embodied metaphors by using larger language models that conceptualise embodies the action of the metaphorical sentence. |
| Outcome: | The proposed model enables interpretation of figurative language when the action of the metaphorical sentence is more embodied. |