Preserving Language Capabilities in Vision-Language Models via Representation Regulation (2026.findings-acl)
Copied to clipboard
| Challenge: | Vision-Language Models (VLMs) provide a unified framework to process both text-only and vision-language tasks. |
| Approach: | They propose a method to reduce the distance between visual and textual representations by introducing a Representation Distribution Difference (RDD) loss. |
| Outcome: | Empirical evidence shows that finetuning VLMs on vision-language data has degraded language capabilities. |
Similar Papers
Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models (2025.findings-acl)
Copied to clipboard
Qin Liu, Chao Shang, Ling Liu, Nikolaos Pappas, Jie Ma, Neha Anna John, Srikanth Doss, Lluis Marquez, Miguel Ballesteros, Yassine Benajiba
| Challenge: | LLaVA-7B demonstrated a decline in safety alignment ability on multi-modal inputs compared to its LLM backbone. |
| Approach: | They propose a method to recover alignment ability from LLM backbone while preserving functional capabilities of VLMs. |
| Outcome: | The proposed framework recovers alignment ability that is inherent in the LLM backbone with minimal impact on fluency and linguistic capabilities of pre-trained VLMs. |
Lost in Embeddings: Information Loss in Vision–Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Experiments reveal connectors substantially distort the local geometry of visual representations, with k-nearest neighbors diverging by 40–60% post-projection, correlating with degradation in retrieval performance. |
| Approach: | They propose two approaches to examine and quantify information loss by analyzing latent representation space. |
| Outcome: | The proposed model improves retrieval performance by analyzing changes in k-nearest neighbor relationships between image representations before and after projection. |
LLMs Can Compensate for Deficiencies in Visual Representations (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a strong language backbone in vision-language models compensates for weak visual features by contextualizing or enriching them. |
| Approach: | They investigate whether strong language backbone compensates for weak visual features . they use CLIP-based vision encoders to perform controlled self-attention ablations . |
| Outcome: | The proposed model compensates for weak visual features by contextualizing or enriching them. |
RWKV-CLIP: A Robust Vision-Language Representation Learner (2024.emnlp-main)
Copied to clipboard
| Challenge: | Using large image-text datasets, large-scale image-data sets have been used for visionlanguage pre-training. |
| Approach: | They propose a framework that leverages Large Language Models to combine and refine information from web-based image-text pairs, synthetic captions, and detection tags. |
| Outcome: | The proposed framework can combine and refine information from web-based image-text pairs, synthetic captions, and detection tags. |
Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction (2024.acl-long)
Copied to clipboard
| Challenge: | EVLGen is a framework for visual-language pre-training with high computational demands. |
| Approach: | They propose a streamlined framework for the pre-training of visually conditioned language generation models with high computational demands. |
| Outcome: | The proposed framework accelerates training of vision-language models by a factor of 5 without compromising performance. |
Red Teaming Visual Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | VLMs (Vision-Language Models) can be induced to generate harmful or inaccurate content through specific test cases. |
| Approach: | They propose a red teaming dataset which encompasses 12 subtasks under 4 primary aspects (faithfulness, privacy, safety, fairness) this dataset is the first to benchmark current VLMs in terms of these 4 aspects . |
| Outcome: | The proposed dataset shows that 10 open-source VLMs struggle with red teaming in different degrees and have up to 31% performance gap with GPT-4V. |
Vision-Language Models Align with Human Neural Representations in Concept Processing (2026.eacl-long)
Copied to clipboard
| Challenge: | Recent studies suggest that transformer-based vision-language models capture the multimodality of concept processing in the human brain. |
| Approach: | They analysed multiple VLMs employing different strategies to integrate visual and textual modalities, along with language-only counterparts. |
| Outcome: | The transformer-based vision-language models outperform language-only models in two experimental conditions, while only some outperformed the language-based models. |
Can VLMs Actually See and Read? A Survey on Modality Collapse in Vision-Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Vision-language models integrate textual and visual information, enabling them to process visual inputs and generate predictions. |
| Approach: | They review work on modality collapse analysis to provide insights into the reason for this unintended behavior and review probing studies for fine-grained vision-language understanding. |
| Outcome: | The proposed models can achieve competitive performance in vision-language tasks despite relying heavily on textual information and ignoring visual information. |
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies show that deep vision-only and language-only models project inputs into a partially aligned representational space. |
| Approach: | They investigate whether a model's representational code is semantically shared . they find that alignment peaks in mid-to-late layers of both model types . |
| Outcome: | a forced-choice "Pick-a-Pic" task shows human preferences for image-caption matches are mirrored in embedding spaces across vision-language model pairs. |
Measuring Progress in Fine-grained Vision-and-Language Understanding (2023.acl-long)
Copied to clipboard
| Challenge: | X-VLM models lack "fine-grained" understanding of relationships, verbs and numbers in images . pretraining on large-scale image–text data from the Web has facilitated rapid progress on many vision-and-language tasks . |
| Approach: | They investigate models that outperform other baselines on fine-grained data . they highlight importance of novel losses and rich data sources for learning fine-grain skills . |
| Outcome: | The proposed model outperforms baseline models on four fine-grained benchmarks . the model outpersforms other baseline models and even degrades performance . |