Challenge: Recent studies have advanced learning VSE under the monolingual setup.
Approach: They propose a model with diverse multi-head attention to learn grounded multilingual multimodal representations by leveraging visual object detection.
Outcome: The proposed model performs well in German-Image and English-Image matching tasks and in the Semantic Textual Similarity task with English descriptions of visual content.

Similar Papers

A Visual Attention Grounding Neural Model for Multimodal Machine Translation (D18-1)

Copied to clipboard

Challenge: Existing approaches to multimodal machine translation do not integrate visual information into the translation process.
Approach: They propose a multimodal machine translation model that utilizes parallel visual and textual information.
Outcome: The proposed model outperforms existing methods on the Multi30K and Ambiguous COCO datasets.
Aligning Multilingual Word Embeddings for Cross-Modal Retrieval Task (D19-64)

Copied to clipboard

Challenge: Existing methods to learn multimodal multilingual embeddings for text and image retrieval tasks are limited to English.
Approach: They propose a new approach to learn multimodal multilingual embeddings for matching images and captions in two languages by combing two existing objective functions and adapting alignment between existing languages.
Outcome: The proposed model achieves state-of-the-art in retrieval and caption-caption tasks while adapting existing language alignments.
Aligning Multilingual Word Embeddings for Cross-Modal Retrieval Task (D19-66)

Copied to clipboard

Challenge: Existing methods to learn multimodal multilingual embeddings for text and image retrieval tasks are limited to English.
Approach: They propose a new approach to learn multimodal multilingual embeddings for matching images and captions in two languages by combing two existing objective functions and adapting alignment between existing languages.
Outcome: The proposed model achieves state-of-the-art in retrieval and caption-caption tasks while adapting existing language alignments.
Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models (2021.naacl-main)

Copied to clipboard

Challenge: a new study examines zero-shot cross-lingual transfer of vision-language models . we study multilingual text-to-video search in non-English languages without annotations .
Approach: They propose a Transformer-based model that learns contextual multilingual multimodal embeddings . they propose 'zero-shot cross-lingual transfer' to improve multilingual search .
Outcome: The proposed model outperforms baselines on multilingual text-to-video search and multilingual image search on VTT and VATEX.
Cross-lingual Cross-modal Pretraining for Multimodal Retrieval (2021.naacl-main)

Copied to clipboard

Challenge: Recent pretrained vision-language models have achieved impressive performance on cross-modal retrieval tasks in English.
Approach: They propose a new approach to learn cross-lingual cross-modal representations for matching images and captions in multiple languages using an annotated corpus.
Outcome: The proposed model achieves impressive performance on two multimodal multilingual image caption benchmarks: Multi30k with German captions and MSCOCO with Japanese captions.
A Multi-task Approach to Learning Multilingual Representations (P18-2)

Copied to clipboard

Challenge: Using a multi-task model, we learn word and sentence embeddings in a single task.
Approach: They propose a multi-task modeling approach that trains a skip-gram model and a cross-lingual sentence similarity model to learn word and sentence embeddings together.
Outcome: The proposed model can learn word and sentence embeddings in a multilingual distributed representations of text using a cross-lingual sentence similarity model.
Cross-lingual Visual Pre-training for Multimodal Machine Translation (2021.eacl-main)

Copied to clipboard

Challenge: Pre-trained language models have been shown to improve performance in many natural language tasks.
Approach: They propose to combine cross-lingual and visual pre-training to learn visually-grounded cross-linguistic representations using masked region classification and three-way parallel vision & language corpora.
Outcome: The proposed models obtain state-of-the-art performance when fine-tuned for multimodal machine translation.
Assessing Multilingual Fairness in Pre-trained Multimodal Representations (2022.findings-acl)

Copied to clipboard

Challenge: Recent pre-trained multimodal models have shown exceptional capabilities towards connecting images and natural language.
Approach: They propose two new fairness notions for pre-trained multimodal models that consider language as the fairness recipient.
Outcome: The proposed models can be generalized to multilingualism by cross-lingual alignment . the results show that the models are individually fair across languages .
Recognizing Multimodal Entailment (2021.acl-tutorials)

Copied to clipboard

Challenge: This tutorial introduces the multimodal entailment task for detecting semantic alignments . the task requires fine-grained understanding of visual and linguistic semantics questions .
Approach: This tutorial introduces the multimodal entailment task to machine learning . it introduces a dataset for recognizing multimodal alignments .
Outcome: This tutorial introduces the multimodal entailment task . it can be useful for detecting semantic alignments when a single modality alone is not enough .
Grounding Multilingual Multimodal LLMs With Cultural Knowledge (2025.emnlp-main)

Copied to clipboard

Challenge: a new data-centric approach could address cultural gaps in multimodal large language models . despite being trained on billions of image-text pairs, today's models are biased towards English and Western data.
Approach: They propose a data-centric approach that directly grounds MLLMs in cultural knowledge.
Outcome: The proposed approach outperforms open-source models on cultural-focused benchmarks without degrading results on mainstream vision–language tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations