Challenge: Large Language Models (LLMs) are trained on large corpora of disembodied texts.
Approach: They propose a multi-modal task of predicting the outcomes of actions solely from realistic sensory inputs (images and text). They extend an LLM to model latent representations of objects to better predict action outcomes in an environment.
Outcome: The proposed model can capture commonsense when augmented with visual information and generalize and learn commonsensical reasoning better.

Similar Papers

What Action Causes This? Towards Naive Physical Action-Effect Prediction (P18-1)

Copied to clipboard

Challenge: a new task on naive physical action-effect prediction addresses the relationship between concrete actions and their effects on the state of the physical world as depicted by images.
Approach: They propose a task that harnesses web image data to facilitate action-effect prediction.
Outcome: The proposed approach harnesses web image data through distant supervision to facilitate learning for action-effect prediction.
Reasoning about Actions and State Changes by Injecting Commonsense Knowledge (D18-1)

Copied to clipboard

Challenge: Recent work has shown impressive progress in comprehending procedural text, but their predictions can be inconsistent or highly improbable.
Approach: They propose to incorporate global constraints and bias reading with corpora-based preferences to improve the predicted effects of actions in a paragraph.
Outcome: The proposed model significantly outperforms earlier models on a benchmark dataset for procedural text comprehension (+8% relative gain) it avoids nonsensical predictions that earlier models make, and it is more robust than previous models.
Making Large Language Models into World Models with Precondition and Effect Knowledge (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are not inherently designed to model real-world dynamics, but can be induced to perform two critical world model functions: determining the applicability of an action based on a given world state and predicting the resulting world state upon action execution.
Approach: They propose to use Large Language Models to model world states and preconditions . they validate that precondition and effect knowledge generated by LLMs aligns with human understanding of world dynamics .
Outcome: The proposed model can predict valid actions and state transitions, thereby replicating existing models.
MORE: Multi-mOdal REtrieval Augmented Generative Commonsense Reasoning (2024.findings-acl)

Copied to clipboard

Challenge: Language Models (LLMs) have gained increasing prominence in artificial intelligence, especially Large Language Model (LLm) due to the well-recognized reporting bias, the recording of commonsense information is significantly less than its existence in reality.
Approach: They propose a Multi-mOdal REtrieval framework to leverage both text and images to enhance commonsense ability of language models.
Outcome: The proposed framework can leverage both text and images to enhance commonsense ability of language models.
A Systematic Investigation of Commonsense Knowledge in Large Language Models (2022.emnlp-main)

Copied to clipboard

Challenge: Recent large language models (LMs) have shown impressive performance on many NLP tasks under the zero-shot and few-shot setup.
Approach: They conduct a systematic and rigorous zero-shot and few-shot commonsense evaluation of large pre-trained language models to better understand their ability to capture commonsensical knowledge.
Outcome: The proposed model can exploit surface cues and annotation artefacts without task-specific supervision and is insufficient to achieve human-level commonsense performance.
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion (2026.acl-short)

Copied to clipboard

Challenge: Large Language Models lack visual grounding on visual reasoning, despite training on text alone.
Approach: They propose a late multi-image fusion method that augments LLMs with test-time visual signals.
Outcome: Using a late multi-image fusion method, the proposed model outperforms LLMs on visual reasoning and matches VLMs in vision-based tasks.
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning (2025.naacl-long)

Copied to clipboard

Challenge: Motivated by in-context learning capabilities of Large Language Models (LLMs), multimodal LLMs with additional visual modality are also exhibited with similar ICL abilities when multiple image-text pairs are provided as demonstrations.
Approach: They conduct systematic and principled evaluation of multimodal ICL for models of different scales on a broad spectrum of new yet critical tasks.
Outcome: The proposed model performance improves on a broad spectrum of new yet critical tasks.
EvEntS ReaLM: Event Reasoning of Entity States via Language Models (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to model event implications fail to reason about the world, despite their knowledge of physical attributes.
Approach: They propose to use a model prompting technique to prompt models of event implications by targeting their understanding of physical attributes.
Outcome: The proposed model prompting technique is especially useful for unseen attributes or when only limited data is available.
Can Textual Unlearning Solve Cross-Modality Safety Alignment? (2024.findings-emnlp)

Copied to clipboard

Challenge: integrating new modalities into large language models creates new attack surface . existing safety training techniques like SFT and RLHF are not feasible in multi-modal settings .
Approach: They explore whether unlearning in the textual domain can be effective for cross-modality safety alignment.
Outcome: The proposed approach reduces the Attack Success Rate (ASR) to less than 8% and preserves the utility.
Can Language Models Understand Physical Concepts? (2023.emnlp-main)

Copied to clipboard

Challenge: Existing language models do not understand basic physical concepts in the human world.
Approach: They propose a method to transfer embodied knowledge from visual models to LMs . they use visual concepts and embodies concepts learned from interaction with the world .
Outcome: The proposed method achieves comparable performance with scaling up parameters of LMs 134.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations