Challenge: Visual scenes often involve multiple people and humans can distinguish between them based on context descriptions about what happened before, their mental/physical states, and intentions.
Approach: They propose a task that tests human-centric commonsense grounding models' ability to distinguish individuals given context descriptions about what happened before and their mental/physical states or intentions.
Outcome: The proposed model outperforms pre-trained and non-pretrained models on 130k commonsense descriptions annotated on 67k images.

Similar Papers

VCD: A Dataset for Visual Commonsense Discovery in Images (2025.findings-acl)

Copied to clipboard

Challenge: Visual commonsense data sets lack visual grounded representations of commonsensense . existing knowledge bases lack visual-based knowledge tied to actual visual scenes .
Approach: They present a large-scale visual commonsense dataset with over 100,000 images and 14 million object-commonsense pairs that integrates both Seen (directly observable) and Unseen (inferrable) commonsens.
Outcome: The proposed model integrates Seen (directly observable) and Unseen (inferrable) commonsense across Property, Action, and Space aspects.
Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for visual grounding rely on the assumption that the given expression must be literal . this impedes the practical deployment of agents in real-world scenarios.
Approach: They propose a visual grounding task that uses intention expressions to locate foreground entities . they build a large-scale IVG dataset with free-form intention expression to promote VG .
Outcome: The proposed method is based on a large-scale intention-driven visual-language (V-L) dataset with free-form intention expressions.
Knowledge Supports Visual Language Grounding: A Case Study on Colour Terms (2020.acl-main)

Copied to clipboard

Challenge: In human cognition, world knowledge supports the perception of object colours . a lot of recent work in Language & Vision has looked at grounding language in real-world sensory information.
Approach: They propose to integrate visual information and object-specific knowledge via hard-coded or learned fusion to improve visual grounding of colour terms in realistic objects.
Outcome: The proposed models outperform a baseline model that predicts colour terms solely from visual inputs but show interesting differences when predicting atypical colours of so-called colour diagnostic objects.
Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Commonsense Norms (2023.emnlp-main)

Copied to clipboard

Challenge: NormLens is a visual-grounded framework for understanding commonsense norms . state-of-the-art models are not well-aligned with human annotation, we show .
Approach: They propose a visual-grounded framework to study commonsense norms by NormLens . they find that models are not well-aligned with human annotation .
Outcome: The proposed model judgments and explanations are not well-aligned with human annotations.
Commonsense Reasoning for Natural Language Processing (2020.acl-tutorials)

Copied to clipboard

Challenge: In this tutorial, we will outline the various types of commonsense knowledge and discuss techniques to gather and represent commonsence knowledge.
Approach: This tutorial will provide researchers with the critical foundations and recent advances in commonsense representation and reasoning.
Outcome: This tutorial will outline the various types of commonsense and discuss techniques to gather and represent commonsence knowledge while highlighting the challenges specific to this type of knowledge (e.g., reporting bias).
Read Before Grounding: Scene Knowledge Visual Grounding via Multi-step Parsing (2025.coling-main)

Copied to clipboard

Challenge: Existing VG datasets use simple textual descriptions with limited attribute and spatial information between images and text.
Approach: They propose a method that transforms visual knowledge into concise, information-dense visual descriptions.
Outcome: The proposed method significantly improves performance of multimodal grounding models.
Visual Commonsense in Pretrained Unimodal and Multimodal Models (2022.naacl-main)

Copied to clipboard

Challenge: Fig. 1 shows how text-only and image-only models can capture commonsense visual attributes, but reporting bias affects their performance.
Approach: They use a Visual Commonsense Tests dataset to validate their findings . they find multimodal models better reconstruct attribute distributions, but are still subject to reporting bias .
Outcome: The proposed model improves on the unimodal and multimodal models, but is still subject to reporting bias.
Understanding ME? Multimodal Evaluation for Fine-grained Visual Commonsense (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models that understand image and text but also cross-reference in-between are lacking in evaluation data resources.
Approach: They propose a multimodal evaluation pipeline to automatically generate question-answer pairs to test models’ understanding of the visual scene, text, and related knowledge.
Outcome: The proposed model can answer the highly semantic VCR question correctly but fails to answer related visual question (Q2), textual question (q3), and background knowledge question ( Q4) as shallow mappings with language priors and unbalanced utilization of information between modalities.
Do GUI Grounders Truly Understand UI Elements? (2026.findings-eacl)

Copied to clipboard

Challenge: Existing grounding models and benchmarks are skewed toward web and mobile environments, neglecting desktop interfaces (especially windows).
Approach: They propose a GUI Grounding Sensitivity Benchmark to assess UI grounding sensitivity to multiple descriptions of the same UI element.
Outcome: The proposed model generates multiple valid instructions per UI element and develops nuanced validation methods to validate them.
Grounded Multimodal Procedural Entity Recognition for Procedural Documents: A New Dataset and Baseline (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to extract procedural knowledge from documents focus on text-only settings, which is insufficient for entity disambiguation.
Approach: They propose a model to detect the entity and the corresponding bounding box groundings in images.
Outcome: The proposed model detects the entity and the corresponding bounding box groundings in image (i.e., visual entities) it is based on a dataset of a WikiHow 1 and EHow 2 document and the results are compared with existing models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations