Papers with REC

8 papers
Prompting Vision-Language Models For Aspect-Controlled Generation of Referring Expressions (2024.findings-naacl)

Copied to clipboard

Challenge: Referring Expression Generation (REG) is the task of generating a descriptive caption that uniquely identifies a given target in the scene.
Approach: They propose an Aspect-Controlled REG task which requires generating a referring expression conditioned on the input aspect(s) by changing the input input such as color, location, action etc.
Outcome: The proposed model beats all prior works in the CIDEr score and achieves comparable performance to training with 100% of real data.
CK-Transformer: Commonsense Knowledge Enhanced Transformers for Referring Expression Comprehension (2023.findings-eacl)

Copied to clipboard

Challenge: Existing frameworks for referring expression comprehension with commonsense knowledge are lacking in the field of multimodal referring .
Approach: They propose a framework for commonsense knowledge Enhanced Transformers which integrates commonsensible knowledge into representations of objects in an image.
Outcome: The proposed framework improves on the existing state of the art in referring expression comprehension with commonsense knowledge (CK-Transformer) it achieves 3.14% accuracy over the existing framework.
MaPPER: Multimodal Prior-guided Parameter Efficient Tuning for Referring Expression Comprehension (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for Referring Expression Comprehension (REC) lack specific domain abilities for precise local visual perception and visual-language alignment.
Approach: They propose a framework for Parameter-Efficient Transfer Learning to localize a visual region via natural language using a prior-guided prior.
Outcome: The proposed framework achieves the best accuracy compared to the current methods with only 1.41% tunable backbone parameters.
Towards Unifying Reference Expression Generation and Comprehension (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models for REG and REC have distinct inputs and connections between them . a new model for REg and reprehension is needed to solve these problems .
Approach: They propose a unified model for REG and REC that fuses image, region and text . they propose Vision-conditioned Masked Language Modeling and Text-Conditioned Region Prediction .
Outcome: The proposed model outperforms existing models on REG and REC tasks.
CLEVR-Implicit: A Diagnostic Dataset for Implicit Reasoning in Referring Expression Comprehension (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that pre-trained vision-language models perform well in cross-modal tasks, including referring expression comprehension.
Approach: They propose a method that enables VL models to reason with implicit text . they propose to use a dataset to align the text with objects in the images .
Outcome: The proposed method improves performance 37.94% on referring expression comprehension task.
FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension (2024.emnlp-main)

Copied to clipboard

Challenge: Referring Expression Comprehension (REC) is a cross-modal task that objectively evaluates the capabilities of language understanding, image comprehension, and language-to-image grounding.
Approach: They propose to use a new reference expression comprehension (REC) dataset to evaluate the capabilities of language understanding, image comprehension, and language-to-image grounding.
Outcome: The proposed model is able to reject scenarios where the target object is not visible in the image, a key aspect often overlooked in existing models and approaches.
RECANTFormer: Referring Expression Comprehension with Varying Numbers of Targets (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for REC assume that a single referring expression always refers to a one instance in the image.
Approach: They propose a one-stage method that generates bounding boxes for objects referred to in natural language expressions.
Outcome: The proposed method outperforms baselines in three GREC datasets.
KnowDR-REC: Auditing Knowledge-Conditioned Visual Grounding in Referring Expression Comprehension (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics suggest that Multimodal large language models have acquired fine-grained visual grounding capabilities.
Approach: They propose a benchmark to assess Referring Expression Comprehension (REC) that uses intra-image visual cues to localize target objects and a controllable evaluation mechanism to test sensitivity to fine-grained factual changes.
Outcome: The proposed benchmarks show that multimodal large language models have a high level of performance on the RefCOCO family of benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations