| Challenge: | Existing visual referring expression recognition tasks have multiple annotation problems and language bias problems. |
| Approach: | They propose a visual grounding task called Visual Query Detection . they evaluate the first algorithms on visual referring expression datasets and VQDv1 datasets . |
| Outcome: | The proposed algorithms are compared with existing visual referring expression comprehension datasets and the new VQDv1 dataset. |
Similar Papers
Uncovering the Full Potential of Visual Grounding Methods in VQA (2024.acl-long)
Copied to clipboard
| Challenge: | Visual Grounding (VG) methods in VQA aim to strengthen a model's reliance on question-relevant visual information. |
| Approach: | They propose to strengthen a model's reliance on question-relevant visual information by using a visual grounding method that is based on a question-related visual input. |
| Outcome: | The proposed methods can be much more effective when evaluation conditions are corrected. |
Read Before Grounding: Scene Knowledge Visual Grounding via Multi-step Parsing (2025.coling-main)
Copied to clipboard
| Challenge: | Existing VG datasets use simple textual descriptions with limited attribute and spatial information between images and text. |
| Approach: | They propose a method that transforms visual knowledge into concise, information-dense visual descriptions. |
| Outcome: | The proposed method significantly improves performance of multimodal grounding models. |
OCID-Ref: A 3D Robotic Dataset With Embodied Language For Clutter Scene Grounding (2021.naacl-main)
Copied to clipboard
| Challenge: | Visual grounding (VG) is a crucial task in natural language processing, computer vision, and robotics. |
| Approach: | They propose a visual grounding task with referring expressions of occluded objects in a OCID-Ref dataset with 2,300 scenes and a point cloud input. |
| Outcome: | The proposed dataset shows that it can handle 2D and 3D signals but referring to occluded objects remains challenging for the modern visual grounding systems. |
ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding (2025.acl-long)
Copied to clipboard
| Challenge: | 3D visual grounding models localize entities in a scene referred to by natural language text . recent studies focused on LLM-based scaling of 3DVG datasets, but these do not capture the full range of potential prompts which could be specified in the English language. |
| Approach: | They propose a framework for linguistically analyzing 3DVG prompts and introduce a diagnostic dataset for evaluating 3D visual grounding methods against a diverse set of language patterns. |
| Outcome: | The proposed framework scales up and tests against a representative set of prompts in the english language. |
CommVQA: Situating Visual Question Answering in Communicative Contexts (2024.emnlp-main)
Copied to clipboard
| Challenge: | Current visual question answering models are trained on image-question pairs in isolation, but the questions people ask are dependent on their informational needs and prior knowledge about the image content. |
| Approach: | They propose a visual question-answer-as-question dataset that contains 1000 images and 8,949 question-announcer pairs to evaluate how situating images within naturalistic contexts shapes visual questions. |
| Outcome: | The proposed dataset contains 1000 images and 8,949 question-answer pairs. |
Find Someone Who: Visual Commonsense Understanding in Human-Centric Grounding (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Visual scenes often involve multiple people and humans can distinguish between them based on context descriptions about what happened before, their mental/physical states, and intentions. |
| Approach: | They propose a task that tests human-centric commonsense grounding models' ability to distinguish individuals given context descriptions about what happened before and their mental/physical states or intentions. |
| Outcome: | The proposed model outperforms pre-trained and non-pretrained models on 130k commonsense descriptions annotated on 67k images. |
Learning to Ground Visual Objects for Visual Dialog (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to ground visual objects are inadequate for visual dialog . a posterior distribution is inferred from context and questions, while posterior distributions are used to facilitate visual objects grounding. |
| Approach: | They propose a method to learn to ground visual objects for visual dialog using prior and posterior distributions over visual objects to facilitate visual objects grounding. |
| Outcome: | The proposed approach improves the existing models in generative and discriminative settings by a significant margin. |
A negative case analysis of visual grounding methods for VQA (2020.acl-main)
Copied to clipboard
| Challenge: | Existing Visual Question Answering (VQA) methods exploit dataset biases and spurious statistical correlations instead of producing correct answers for the right reasons. |
| Approach: | They propose to incorporate visual cues to better ground VQA models . they also propose a regularization effect which prevents over-fitting to linguistic priors . |
| Outcome: | The proposed method outperforms existing methods on the Visual Question Answering (VQA) dataset. |
Visual Referring Expression Recognition: What Do Systems Actually Learn? (N18-2)
Copied to clipboard
| Challenge: | Existing systems for referring expression recognition ignore linguistic structure, instead relying on shallow correlations introduced by unintended biases in the data selection and annotation process. |
| Approach: | They propose to use a system trained on the input image without the input referring expression to achieve a precision of 71.2% in top-2 predictions. |
| Outcome: | The proposed model can achieve 71.2% accuracy on the input image without the input referring expression and 84.2% on the object category given the input. |
Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for visual grounding rely on the assumption that the given expression must be literal . this impedes the practical deployment of agents in real-world scenarios. |
| Approach: | They propose a visual grounding task that uses intention expressions to locate foreground entities . they build a large-scale IVG dataset with free-form intention expression to promote VG . |
| Outcome: | The proposed method is based on a large-scale intention-driven visual-language (V-L) dataset with free-form intention expressions. |