Challenge: Despite recent progress towards scaling up multimodal vision-language models, these models struggle on compositional generalization benchmarks such as Winoground.
Approach: They propose to use a cross-modal attention regularization loss to enforce relation alignment by capturing the semantic relation ‘in’ to match the visual attention from the mug to the grass.
Outcome: The proposed approach improves Winoground Group score by 5.75 points .

Similar Papers

Learning Relation Alignment for Calibrated Cross-modal Retrieval (2021.acl-long)

Copied to clipboard

Challenge: despite advances in multimodal pre-training, cross-modal retrieval remains challenging . lack of relation consistency impairs contextualized representation of image-text pairs .
Approach: They propose a new metric to quantify the relation consistency by measuring the semantic distance between linguistic and visual relations.
Outcome: The proposed method boosts the performance of prevailing models on Flickr30k and MS COCO datasets by a considerable margin.
Modality Alignment between Deep Representations for Effective Video-and-Language Learning (2022.lrec-1)

Copied to clipboard

Challenge: Existing Video-and-Language models do not take into account the different characteristics of video and text representations.
Approach: They propose a method that exploits Centered Kernel Alignment (CKA) to enhance cross-modality attention by combining multiple modalities.
Outcome: The proposed method outperforms conventional multi-modal methods significantly on video QA tasks with +3.57% accuracy increment compared to the baseline in a popular benchmark dataset.
Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models (2026.findings-acl)

Copied to clipboard

Challenge: a recent study shows that vision-language models have modality gaps that persist even in well-aligned models.
Approach: They propose a modality-dominance score to measure and leverage modality gaps . they propose automatic interpretability metrics to evaluate these features in a scalable manner .
Outcome: The proposed framework allows for training-free probing and editing methods for understanding model perception across genders and generating adversarial examples.
Attention as Grounding: Exploring Textual and Cross-Modal Attention on Entities and Relations in Language-and-Vision Transformer (2022.findings-acl)

Copied to clipboard

Challenge: Existing work has focused on what is captured by multi-modal architectures.
Approach: They propose a multi-modal transformer that learns syntactic and semantic representations about entities and relations grounded in objects at the level of masked self-attention and cross-modal attention.
Outcome: The proposed model learns syntactic and semantic representations about objects and relations cross-modally and unimodally.
Improving Cross-Modal Alignment in Vision Language Navigation via Syntactic Information (2021.naacl-main)

Copied to clipboard

Challenge: Existing vision language navigation tasks require soft attention over words to locate instructions . a new approach uses syntax information to ground instructions with visual information .
Approach: They propose a vision language navigation agent that utilizes syntax information to enhance alignment between the instruction and the current visual scenes.
Outcome: The proposed agent outperforms the baseline model that does not use syntax information on the Room-to-Room dataset, especially in the unseen environment.
VLA-Mark: A cross modal watermark for large vision-language alignment models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing text watermarking methods disrupt visual-textual alignment, leaving semantic-critical concepts vulnerable.
Approach: They propose a vision-aligned framework that embeds detectable watermarks into outputs . they combine localized patch affinity, global semantic coherence, contextual attention patterns .
Outcome: The proposed framework shows lower PPL and higher BLEU than conventional methods with near-perfect detection (98.8% AUC).
ITA: Image-Text Alignments for Multi-Modal Named Entity Recognition (2022.naacl-main)

Copied to clipboard

Challenge: Recent work on Multi-modal Named Entity Recognition (MNER) relies on image information to model interactions between image and text representations.
Approach: They propose to align image features into the textual space to better utilize attention mechanisms . they use regional object tags, captions and optical characters as visual contexts .
Outcome: The proposed model can achieve state-of-the-art accuracy on multi-modal Named Entity Recognition datasets even without image information.
Low-resource Neural Machine Translation with Cross-modal Alignment (2022.emnlp-main)

Copied to clipboard

Challenge: Existing neural machine translation techniques rely on large monolingual corpus, which is costly for some low-resource languages.
Approach: They propose a cross-modal contrastive learning method to learn a shared space for all languages by additional visual modality.
Outcome: The proposed method can learn cross-modal and cross-lingual alignment with small amount of image-text pairs and achieves significant improvements over the text-only baseline.
Understanding Attention for Vision-and-Language Tasks (2022.coling-1)

Copied to clipboard

Challenge: Attention mechanism has been used in Vision-and-Language (VL) tasks to bridge the semantic gap between visual and textual clues.
Approach: They conduct a comprehensive analysis on understanding the role of attention alignment by looking into attention score calculation methods and checking how it represents the visual region’s and textual token’s significance for the global assessment.
Outcome: The attention score calculation methods represent visual region’s and textual token’s significance for the global assessment.
Probing Relative Interaction and Dynamic Calibration in Multi-modal Entity Alignment (2025.acl-long)

Copied to clipboard

Challenge: Current methods for multi-modal entity alignment ignore relative interactions between modalities and the accuracy of weights.
Approach: They propose a relative interaction and calibration framework for multi-modal entity alignment that uses attention mechanisms to perceive the uncertainty of the weight for each modality.
Outcome: The proposed framework outperforms baselines across 5 datasets and 23 settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations