Papers with VLU
Improving the Efficiency of Visually Augmented Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Autoregressive Language Models lack visual knowledge due to reporting bias in textual corpora. |
| Approach: | They propose to use visual representations obtained from CLIP multimodal system to augment autoregressive language models with visual knowledge. |
| Outcome: | The proposed model outperforms VALM for visual language understanding, natural language understanding and language modeling tasks despite being significantly more efficient and simpler. |
XtremeCLIP: Extremely Parameter-efficient Tuning for Low-resource Vision Language Understanding (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to fine-tune visual-language understanding (VLU) require tasks-specific designs and sufficient training data. |
| Approach: | They propose a simple yet efficient paradigm for low-resource Visual Language Understanding (VLU) they reformulate a series of VLU tasks as an open-book affinity-matching problem. |
| Outcome: | The proposed framework outperforms baselines in low-resource settings. |
Distilled Dual-Encoder Model for Vision-Language Understanding (2022.emnlp-main)
Copied to clipboard
| Challenge: | Experimental results show that the proposed cross-modal attention distillation is crucial to the success of our framework. |
| Approach: | They propose a framework that distills knowledge of fusion-encoder teacher into dual-encoding student model. |
| Outcome: | The proposed model is competitive with the fusion-encoder teacher model in performance, but suffers from a lack of deep cross-modal interactions. |
SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language Understanding (2026.acl-long)
Copied to clipboard
| Challenge: | Existing block-wise discrete diffusion models lack robust autoregressive (AR) decoders. |
| Approach: | They propose a block-wise discrete diffusion framework for large-scale vision-language understanding with a progressive beta noise curriculum. |
| Outcome: | The proposed framework improves training efficiency, convergence stability, and task performance over conventional block diffusion. |