Papers with REC
Prompting Vision-Language Models For Aspect-Controlled Generation of Referring Expressions (2024.findings-naacl)
Copied to clipboard
Danfeng Guo, Sanchit Agarwal, Arpit Gupta, Jiun-Yu Kao, Emre Barut, Tagyoung Chung, Jing Huang, Mohit Bansal
| Challenge: | Referring Expression Generation (REG) is the task of generating a descriptive caption that uniquely identifies a given target in the scene. |
| Approach: | They propose an Aspect-Controlled REG task which requires generating a referring expression conditioned on the input aspect(s) by changing the input input such as color, location, action etc. |
| Outcome: | The proposed model beats all prior works in the CIDEr score and achieves comparable performance to training with 100% of real data. |
CK-Transformer: Commonsense Knowledge Enhanced Transformers for Referring Expression Comprehension (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing frameworks for referring expression comprehension with commonsense knowledge are lacking in the field of multimodal referring . |
| Approach: | They propose a framework for commonsense knowledge Enhanced Transformers which integrates commonsensible knowledge into representations of objects in an image. |
| Outcome: | The proposed framework improves on the existing state of the art in referring expression comprehension with commonsense knowledge (CK-Transformer) it achieves 3.14% accuracy over the existing framework. |
MaPPER: Multimodal Prior-guided Parameter Efficient Tuning for Referring Expression Comprehension (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for Referring Expression Comprehension (REC) lack specific domain abilities for precise local visual perception and visual-language alignment. |
| Approach: | They propose a framework for Parameter-Efficient Transfer Learning to localize a visual region via natural language using a prior-guided prior. |
| Outcome: | The proposed framework achieves the best accuracy compared to the current methods with only 1.41% tunable backbone parameters. |
Towards Unifying Reference Expression Generation and Comprehension (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing models for REG and REC have distinct inputs and connections between them . a new model for REg and reprehension is needed to solve these problems . |
| Approach: | They propose a unified model for REG and REC that fuses image, region and text . they propose Vision-conditioned Masked Language Modeling and Text-Conditioned Region Prediction . |
| Outcome: | The proposed model outperforms existing models on REG and REC tasks. |
CLEVR-Implicit: A Diagnostic Dataset for Implicit Reasoning in Referring Expression Comprehension (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies show that pre-trained vision-language models perform well in cross-modal tasks, including referring expression comprehension. |
| Approach: | They propose a method that enables VL models to reason with implicit text . they propose to use a dataset to align the text with objects in the images . |
| Outcome: | The proposed method improves performance 37.94% on referring expression comprehension task. |
FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension (2024.emnlp-main)
Copied to clipboard
| Challenge: | Referring Expression Comprehension (REC) is a cross-modal task that objectively evaluates the capabilities of language understanding, image comprehension, and language-to-image grounding. |
| Approach: | They propose to use a new reference expression comprehension (REC) dataset to evaluate the capabilities of language understanding, image comprehension, and language-to-image grounding. |
| Outcome: | The proposed model is able to reject scenarios where the target object is not visible in the image, a key aspect often overlooked in existing models and approaches. |
RECANTFormer: Referring Expression Comprehension with Varying Numbers of Targets (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for REC assume that a single referring expression always refers to a one instance in the image. |
| Approach: | They propose a one-stage method that generates bounding boxes for objects referred to in natural language expressions. |
| Outcome: | The proposed method outperforms baselines in three GREC datasets. |
KnowDR-REC: Auditing Knowledge-Conditioned Visual Grounding in Referring Expression Comprehension (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation metrics suggest that Multimodal large language models have acquired fine-grained visual grounding capabilities. |
| Approach: | They propose a benchmark to assess Referring Expression Comprehension (REC) that uses intra-image visual cues to localize target objects and a controllable evaluation mechanism to test sensitivity to fine-grained factual changes. |
| Outcome: | The proposed benchmarks show that multimodal large language models have a high level of performance on the RefCOCO family of benchmarks. |