Can Language Models Understand Physical Concepts? (2023.emnlp-main)

Copied to clipboard

Challenge: Existing language models do not understand basic physical concepts in the human world.
Approach: They propose a method to transfer embodied knowledge from visual models to LMs . they use visual concepts and embodies concepts learned from interaction with the world .
Outcome: The proposed method achieves comparable performance with scaling up parameters of LMs 134.

Similar Papers

Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models? (2025.findings-emnlp)

Copied to clipboard

Challenge: Despite significant progress in multimodal language models, it remains unclear whether visual grounding enhances their understanding of embodied knowledge compared to text-only models.
Approach: They propose to assess vision-language models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions.
Outcome: The proposed benchmark assesses the models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions.
POSQA: Probe the World Models of LLMs with Size Comparisons (2023.findings-emnlp)

Copied to clipboard

Challenge: Embodied language comprehension emphasizes that language understanding is not only mental processing in the brain but also involves interactions with the physical and social environment.
Approach: They propose to use a physical object size question to examine the extremity of large language models to test their embodied comprehension.
Outcome: The proposed dataset shows that even the largest LLMs perform poorly under the zero-shot setting.
NEWTON: Are Large Language Models Capable of Physical Reasoning? (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models have been shown to encapsulate syntactic, semantic, word sense, and common-sense knowledge, but limited exploration of their physical reasoning abilities has been conducted.
Approach: They propose a repository and benchmark to evaluate LLMs' physical reasoning skills . they use a pipeline to generate a variant of the benchmark customized to the objects and attributes relevant for their application.
Outcome: The proposed benchmark examines the reasoning capabilities of language models across reasoning tasks.
VIPHY: Probing “Visible” Physical Commonsense Knowledge (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have demonstrated that vision-language models can retain and generalize knowledge, but they do not measure their ability to retain it.
Approach: They build an automatic pipeline to derive a knowledge resource for calibrating and probing vision-language models.
Outcome: The proposed model outperforms the pretrained model on size and spatial tasks.
Penetrative AI: Making LLMs Comprehend the Physical World (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities across a range of tasks.
Approach: They explore how LLMs can be extended to interact with and reason about the physical world through IoT sensors and actuators, a concept that they call "Penetrative AI".
Outcome: The proposed approach extends LLMs' capabilities to interact with and reason about the physical world through IoT sensors and actuators.
Exploring Spatial Schema Intuitions in Large Language and Vision Models (2024.findings-acl)

Copied to clipboard

Challenge: Large language models excel in varied NLP tasks, but lack a direct connection between sensory perception and physical action.
Approach: They examine whether large language models capture implicit human intuitions about building blocks of language . they employ spatial cognitive foundations developed through early sensorimotor experiences .
Outcome: The proposed model captures implicit human intuitions about building blocks of language without a tangible connection to embodied experiences.
Language Models Don’t Learn the Physical Manifestation of Language (2024.acl-long)

Copied to clipboard

Challenge: a new study examines the differences between language-only models and humans . we show that language-based models do not learn the physical manifestation of language .
Approach: They propose to use a series of tasks to investigate visual-auditory properties of language to test their hypothesis.
Outcome: The proposed model does not learn the physical manifestation of language . the results highlight the limitations of linguistic knowledge acquired without sensory experience .
Embodied Language Learning: Opportunities, Challenges, and Future Directions (2024.findings-acl)

Copied to clipboard

Challenge: embodied language learning is a form of language understanding where the language learner is situated in the world, perceives it, and interacts with it.
Approach: They propose to use a concept of World Scopes to measure progress in language understanding research.
Outcome: The proposed framework identifies gaps and suggests future directions for language understanding research.
SemVink: Advancing VLMs’ Semantic Understanding of Optical Illusions via Visual Global Thinking (2025.emnlp-main)

Copied to clipboard

Challenge: Vision-language models excel in semantic tasks but fail at detecting hidden content . current architectures prioritize abstract reasoning over low-level visual operations .
Approach: They propose a benchmark to test vision-language models that can detect hidden content . they propose HC-Bench to scale images to low resolutions to unlock 99% accuracy .
Outcome: HC-Bench shows that leading VLMs achieve near-zero accuracy even with explicit prompting . et al.: current models prioritize abstract reasoning over low-level visual operations . they urge a shift toward hybrid models bridging gap between computational vision and human cognition .
LMs stand their Ground: Investigating the Effect of Embodiment in Figurative Language Interpretation by Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Figures are based on the use of words in a way that deviates from their conventional order and meaning.
Approach: They propose to use a figurative language model to interpret embodied metaphors by using larger language models that conceptualise embodies the action of the metaphorical sentence.
Outcome: The proposed model enables interpretation of figurative language when the action of the metaphorical sentence is more embodied.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations