Challenge: Visual referring expression comprehension (ReC) models can be trained for a domain, but it remains unclear if they can be applied in a zero-shot manner to more complex tasks like ReC.
Approach: They propose a method that repurposes CLIP, a state-of-the-art large-scale model, for training a referring expression comprehension model for a new visual domain.
Outcome: The proposed model reduces the gap between zero-shot baselines from prior work and supervised models by as much as 29% on RefCOCOg, and on ReFGTA (video game imagery), and its relative improvement over supervised ReC models is 8%.

Similar Papers

FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension (2024.emnlp-main)

Copied to clipboard

Challenge: Referring Expression Comprehension (REC) is a cross-modal task that objectively evaluates the capabilities of language understanding, image comprehension, and language-to-image grounding.
Approach: They propose to use a new reference expression comprehension (REC) dataset to evaluate the capabilities of language understanding, image comprehension, and language-to-image grounding.
Outcome: The proposed model is able to reject scenarios where the target object is not visible in the image, a key aspect often overlooked in existing models and approaches.
Text Augmented Spatial Aware Zero-shot Referring Image Segmentation (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing zero-shot referring image segmentation methods focus on global-level alignment of image-text pairs, neglecting fine-grained matching between referring sentence and local image regions.
Approach: They propose a zero-shot referring image segmentation task that is training-free . they use a mask proposal network and a text-augmented spatial-correction score .
Outcome: The proposed method outperforms state-of-the-art zero-shot referring image segmentation methods.
KnowDR-REC: Auditing Knowledge-Conditioned Visual Grounding in Referring Expression Comprehension (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics suggest that Multimodal large language models have acquired fine-grained visual grounding capabilities.
Approach: They propose a benchmark to assess Referring Expression Comprehension (REC) that uses intra-image visual cues to localize target objects and a controllable evaluation mechanism to test sensitivity to fine-grained factual changes.
Outcome: The proposed benchmarks show that multimodal large language models have a high level of performance on the RefCOCO family of benchmarks.
RECANTFormer: Referring Expression Comprehension with Varying Numbers of Targets (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for REC assume that a single referring expression always refers to a one instance in the image.
Approach: They propose a one-stage method that generates bounding boxes for objects referred to in natural language expressions.
Outcome: The proposed method outperforms baselines in three GREC datasets.
TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to enhance zero-shot abilities in image captioning fail with fine-grained datasets.
Approach: They propose a method to enhance captions with additional object-part details using object detector proposals and natural language processing techniques.
Outcome: The proposed method improves performance on fine-grained datasets and improves on existing methods.
Scene Graph Enhanced Pseudo-Labeling for Referring Expression Comprehension (2023.findings-emnlp)

Copied to clipboard

Challenge: Referring expression comprehension is a visual-linguistic task that involves localizing objects in images based on textual referring expressions.
Approach: They propose a scene graph-based framework that generates high-quality pseudo region-query pairs . their method captures relationships between objects in images and generates expressions enriched with relation information.
Outcome: The proposed framework outperforms existing methods by 10%, 12%, and 11% on RefCOCO, RefCoCO+, and Ref COCOg datasets.
Towards Unifying Reference Expression Generation and Comprehension (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models for REG and REC have distinct inputs and connections between them . a new model for REg and reprehension is needed to solve these problems .
Approach: They propose a unified model for REG and REC that fuses image, region and text . they propose Vision-conditioned Masked Language Modeling and Text-Conditioned Region Prediction .
Outcome: The proposed model outperforms existing models on REG and REC tasks.
Prompting Vision-Language Models For Aspect-Controlled Generation of Referring Expressions (2024.findings-naacl)

Copied to clipboard

Challenge: Referring Expression Generation (REG) is the task of generating a descriptive caption that uniquely identifies a given target in the scene.
Approach: They propose an Aspect-Controlled REG task which requires generating a referring expression conditioned on the input aspect(s) by changing the input input such as color, location, action etc.
Outcome: The proposed model beats all prior works in the CIDEr score and achieves comparable performance to training with 100% of real data.
Zero-Shot Entity Linking by Reading Entity Descriptions (P19-1)

Copied to clipboard

Challenge: Existing approaches to link entities to unseen entities require in-domain labeled data.
Approach: They propose a zero-shot entity linking task where mentions must be linked to unseen entities without in-domain labeled data.
Outcome: The proposed task can generalize to unseen entities without metadata or alias tables . the proposed system improves over baselines, including BERT, on a new dataset .
Visual Referring Expression Recognition: What Do Systems Actually Learn? (N18-2)

Copied to clipboard

Challenge: Existing systems for referring expression recognition ignore linguistic structure, instead relying on shallow correlations introduced by unintended biases in the data selection and annotation process.
Approach: They propose to use a system trained on the input image without the input referring expression to achieve a precision of 71.2% in top-2 predictions.
Outcome: The proposed model can achieve 71.2% accuracy on the input image without the input referring expression and 84.2% on the object category given the input.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations