Challenge: Existing fine-tuning approaches for compositional understanding compromise performance in zero-shot multi-modal tasks.
Approach: They propose a method to enhance compositional understanding in pre-trained vision and language models without sacrificing performance in zero-shot multi-modal tasks.
Outcome: The proposed method achieves compositionality on par with state-of-the-art models and retains strong multi-modal capabilities.

Similar Papers

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality (2026.acl-long)

Copied to clipboard

Challenge: a contrastive learning approach for vision-language models is needed to capture compositional information.
Approach: They propose a framework that masks compositional concepts in one modality and reconstructs them conditioned on full contextual information from the other .
Outcome: The proposed framework enhances compositionality in visual language models and improves their ability to capture syntactic structure and linguistic information.
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder (2025.acl-long)

Copied to clipboard

Challenge: Recent studies show that CLIP models struggle with visual reasoning tasks . despite the success of Contrastive Language-Image Pretraining, there are still limitations .
Approach: They propose to use a visual encoder to train CLIP-like models for fine-grained visual reasoning tasks.
Outcome: The proposed models outperform CLIP-like encoders in visual reasoning tasks . the study highlights the importance of VLM architectural choices .
Does CLIP Bind Concepts? Probing Compositionality in Large Image Models (2024.findings-eacl)

Copied to clipboard

Challenge: Large-scale neural network models combining text and images have made incredible progress in recent years, but to what extent they encode compositional representations of the concepts over which they operate remains an open question .
Approach: They compare the performance of a large pretrained vision and language model (CLIP) to a set of three synthetic datasets designed to test concept binding.
Outcome: The proposed model can encode compositional concepts and bind variables in a structure-sensitive way, e.g., differentiating ‘cube behind sphere’ from ‘cub behind cube’.
Vision-Language Model Fine-Tuning via Simple Parameter-Efficient Modification (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in fine-tuning Vision-Language Models have seen the success of prompt tuning and adapter tuning.
Approach: They propose a method to fine-tune CLIP without introducing any overhead of extra parameters.
Outcome: The proposed method improves CLIP by 7.27% average harmonic mean accuracy.
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in multimodal systems have demonstrated remarkable capabilities in generating multimodal content from multimodal inputs.
Approach: They propose a benchmark that leverages large language models to generate deceptive text samples to exploit compositional vulnerabilities across different modalities.
Outcome: The proposed approach exploits compositional vulnerabilities across images, videos, and audios.
Text encoders bottleneck compositionality in contrastive vision-language models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing multimodal models are often unable to reason about simple spatial relations or attribute attachments.
Approach: They first curate CompPrompts, a set of increasingly compositional image captions that VL models should be able to capture . then train text-only recovery probes that aim to reconstruct captions from single-vector text representations produced by several VL model.
Outcome: The proposed model can reconstruct captions from single-vector text representations produced by several models on a broader range of scenes compared to previous models.
CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual Entailment (2022.acl-long)

Copied to clipboard

Challenge: Previously, CLIP was only regarded as a powerful visual encoder.
Approach: They propose a parameter-efficient fine-tuning strategy to boost CLIP's few-shot performance on a visual entailment task without introducing any additional pre-training procedure.
Outcome: The proposed strategy achieves competitive zero/few-shot results on visual question answering and visual entailment tasks without introducing any additional pre-training procedure.
LLMs Can Compensate for Deficiencies in Visual Representations (2025.findings-emnlp)

Copied to clipboard

Challenge: a strong language backbone in vision-language models compensates for weak visual features by contextualizing or enriching them.
Approach: They investigate whether strong language backbone compensates for weak visual features . they use CLIP-based vision encoders to perform controlled self-attention ablations .
Outcome: The proposed model compensates for weak visual features by contextualizing or enriching them.
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for vision-language models treat compositionality and long-caption understanding in isolation.
Approach: They analyze when compositional reasoning and long-caption understanding transfer across tasks and when this relationship fails.
Outcome: The proposed model can generalize on poorly grounded captions and with strong visual grounding, while architectural choices can limit compositional learning.
Can Textual Unlearning Solve Cross-Modality Safety Alignment? (2024.findings-emnlp)

Copied to clipboard

Challenge: integrating new modalities into large language models creates new attack surface . existing safety training techniques like SFT and RLHF are not feasible in multi-modal settings .
Approach: They explore whether unlearning in the textual domain can be effective for cross-modality safety alignment.
Outcome: The proposed approach reduces the Attack Success Rate (ASR) to less than 8% and preserves the utility.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations