VQD: Visual Query Detection In Natural Scenes (N19-1)

Copied to clipboard

Challenge: Existing visual referring expression recognition tasks have multiple annotation problems and language bias problems.
Approach: They propose a visual grounding task called Visual Query Detection . they evaluate the first algorithms on visual referring expression datasets and VQDv1 datasets .
Outcome: The proposed algorithms are compared with existing visual referring expression comprehension datasets and the new VQDv1 dataset.

Similar Papers

Uncovering the Full Potential of Visual Grounding Methods in VQA (2024.acl-long)

Copied to clipboard

Challenge: Visual Grounding (VG) methods in VQA aim to strengthen a model's reliance on question-relevant visual information.
Approach: They propose to strengthen a model's reliance on question-relevant visual information by using a visual grounding method that is based on a question-related visual input.
Outcome: The proposed methods can be much more effective when evaluation conditions are corrected.
Read Before Grounding: Scene Knowledge Visual Grounding via Multi-step Parsing (2025.coling-main)

Copied to clipboard

Challenge: Existing VG datasets use simple textual descriptions with limited attribute and spatial information between images and text.
Approach: They propose a method that transforms visual knowledge into concise, information-dense visual descriptions.
Outcome: The proposed method significantly improves performance of multimodal grounding models.
OCID-Ref: A 3D Robotic Dataset With Embodied Language For Clutter Scene Grounding (2021.naacl-main)

Copied to clipboard

Challenge: Visual grounding (VG) is a crucial task in natural language processing, computer vision, and robotics.
Approach: They propose a visual grounding task with referring expressions of occluded objects in a OCID-Ref dataset with 2,300 scenes and a point cloud input.
Outcome: The proposed dataset shows that it can handle 2D and 3D signals but referring to occluded objects remains challenging for the modern visual grounding systems.
ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding (2025.acl-long)

Copied to clipboard

Challenge: 3D visual grounding models localize entities in a scene referred to by natural language text . recent studies focused on LLM-based scaling of 3DVG datasets, but these do not capture the full range of potential prompts which could be specified in the English language.
Approach: They propose a framework for linguistically analyzing 3DVG prompts and introduce a diagnostic dataset for evaluating 3D visual grounding methods against a diverse set of language patterns.
Outcome: The proposed framework scales up and tests against a representative set of prompts in the english language.
CommVQA: Situating Visual Question Answering in Communicative Contexts (2024.emnlp-main)

Copied to clipboard

Challenge: Current visual question answering models are trained on image-question pairs in isolation, but the questions people ask are dependent on their informational needs and prior knowledge about the image content.
Approach: They propose a visual question-answer-as-question dataset that contains 1000 images and 8,949 question-announcer pairs to evaluate how situating images within naturalistic contexts shapes visual questions.
Outcome: The proposed dataset contains 1000 images and 8,949 question-answer pairs.
Find Someone Who: Visual Commonsense Understanding in Human-Centric Grounding (2022.findings-emnlp)

Copied to clipboard

Challenge: Visual scenes often involve multiple people and humans can distinguish between them based on context descriptions about what happened before, their mental/physical states, and intentions.
Approach: They propose a task that tests human-centric commonsense grounding models' ability to distinguish individuals given context descriptions about what happened before and their mental/physical states or intentions.
Outcome: The proposed model outperforms pre-trained and non-pretrained models on 130k commonsense descriptions annotated on 67k images.
Learning to Ground Visual Objects for Visual Dialog (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to ground visual objects are inadequate for visual dialog . a posterior distribution is inferred from context and questions, while posterior distributions are used to facilitate visual objects grounding.
Approach: They propose a method to learn to ground visual objects for visual dialog using prior and posterior distributions over visual objects to facilitate visual objects grounding.
Outcome: The proposed approach improves the existing models in generative and discriminative settings by a significant margin.
A negative case analysis of visual grounding methods for VQA (2020.acl-main)

Copied to clipboard

Challenge: Existing Visual Question Answering (VQA) methods exploit dataset biases and spurious statistical correlations instead of producing correct answers for the right reasons.
Approach: They propose to incorporate visual cues to better ground VQA models . they also propose a regularization effect which prevents over-fitting to linguistic priors .
Outcome: The proposed method outperforms existing methods on the Visual Question Answering (VQA) dataset.
Visual Referring Expression Recognition: What Do Systems Actually Learn? (N18-2)

Copied to clipboard

Challenge: Existing systems for referring expression recognition ignore linguistic structure, instead relying on shallow correlations introduced by unintended biases in the data selection and annotation process.
Approach: They propose to use a system trained on the input image without the input referring expression to achieve a precision of 71.2% in top-2 predictions.
Outcome: The proposed model can achieve 71.2% accuracy on the input image without the input referring expression and 84.2% on the object category given the input.
Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for visual grounding rely on the assumption that the given expression must be literal . this impedes the practical deployment of agents in real-world scenarios.
Approach: They propose a visual grounding task that uses intention expressions to locate foreground entities . they build a large-scale IVG dataset with free-form intention expression to promote VG .
Outcome: The proposed method is based on a large-scale intention-driven visual-language (V-L) dataset with free-form intention expressions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations