Challenge: Recent studies show that pre-trained vision-language models perform well in cross-modal tasks, including referring expression comprehension.
Approach: They propose a method that enables VL models to reason with implicit text . they propose to use a dataset to align the text with objects in the images .
Outcome: The proposed method improves performance 37.94% on referring expression comprehension task.

Similar Papers

PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language Models (2022.emnlp-main)

Copied to clipboard

Challenge: Recent advances on self-supervised learning have led to powerful vision-language pre-training models that achieve state-of-the-art performance on a wide range of cross-modal tasks.
Approach: They propose a vision-language pre-training framework that reformulates discretized object positions and language in a unified language modeling framework.
Outcome: The proposed model improves performance on position-sensitive vision-language (VL) tasks and also improves on position insensitive tasks.
Conditioned Masked Language and Image Modeling for Image-Text Dense Retrieval (2022.findings-emnlp)

Copied to clipboard

Challenge: Large-scale two-stream pre-trained models like CLIP have achieved tremendous success in image-text retrieval.
Approach: They propose a cross-modal framework for image-text retrieval using two-stream pre-trained models . they embed images and texts into instance representations with two separate encoders . experimental results on MSCOCO and Flickr30k reveal the effectiveness of their framework .
Outcome: The proposed framework improves image-text retrieval performance on two popular cross-modal retrieval benchmarks.
Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for vision-language pre-training lack high-level semantics and text is not sufficiently involved in masked modeling.
Approach: They propose a semantics-enhanced cross-modal MIM framework for vision-language representation learning that harvests high-level semantics from global image features via self-supervised agreement learning and transfers them to local patch encodings by sharing the encode space.
Outcome: The proposed model achieves state-of-the-art or competitive performance on multiple vision-language tasks.
FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension (2024.emnlp-main)

Copied to clipboard

Challenge: Referring Expression Comprehension (REC) is a cross-modal task that objectively evaluates the capabilities of language understanding, image comprehension, and language-to-image grounding.
Approach: They propose to use a new reference expression comprehension (REC) dataset to evaluate the capabilities of language understanding, image comprehension, and language-to-image grounding.
Outcome: The proposed model is able to reject scenarios where the target object is not visible in the image, a key aspect often overlooked in existing models and approaches.
Discovering the Unknown Knowns: Turning Implicit Knowledge in the Dataset into Explicit Training Examples for Visual Question Answering (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods address this issue by introducing an auxiliary task such as visual grounding, cycle consistency, or debiasing.
Approach: They propose a data augmentation pipeline to turn “known” knowledge into training examples for VQA.
Outcome: The proposed model can handle multi-modal information and is based on human-annotated examples.
ConnPrompt: Connective-cloze Prompt Learning for Implicit Discourse Relation Recognition (2022.coling-1)

Copied to clipboard

Challenge: Existing paradigms for Implicit Discourse Relation Recognition (IDRR) do not exploit linguistic evidence embedded in the pre-training process.
Approach: They propose a new paradigm to detect and classify relation sense between two text segments without an explicit connective.
Outcome: The proposed method significantly outperforms the state-of-the-art algorithms even with fewer training data.
Stop Pre-Training: Adapt Visual-Language Models to Unseen Languages (2023.acl-short)

Copied to clipboard

Challenge: Existing studies have shown that the pre-training in English does not transfer well to other languages in a zero-shot setting.
Approach: They propose a simple yet efficient approach to adapt VLP to unseen languages using MPLM.
Outcome: The proposed approach outperforms state-of-the-art models without large parallel corpora across three tasks.
Unsupervised Vision-and-Language Pre-training Without Parallel Images and Captions (2021.naacl-main)

Copied to clipboard

Challenge: Existing models require large amounts of image-caption data for pre-training . existing models require expensive data collection and curation .
Approach: They propose to conduct "mask-and-predict" pre-training on text-only and image-only corpora and introduce the object tags detected by an object recognition model as anchor points to bridge two modalities.
Outcome: The proposed approach achieves performance close to a model pre-trained with aligned data, on four English benchmarks.
Towards Unifying Reference Expression Generation and Comprehension (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models for REG and REC have distinct inputs and connections between them . a new model for REg and reprehension is needed to solve these problems .
Approach: They propose a unified model for REG and REC that fuses image, region and text . they propose Vision-conditioned Masked Language Modeling and Text-Conditioned Region Prediction .
Outcome: The proposed model outperforms existing models on REG and REC tasks.
Visually-Enhanced Phrase Understanding (2023.findings-acl)

Copied to clipboard

Challenge: Large-scale vision-language pre-training models generate high-quality textual representations, which often outperform models that are purely text-based, such as BERT.
Approach: They propose to utilize both textual and visual encoders of multi-modal pre-trained models to enhance language understanding tasks by generating an image associated with a textual prompt.
Outcome: The proposed method outperforms models that are purely text-based on visual and textual understanding tasks and significantly improves the entity clustering task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations