Multi-Head Attention with Diversity for Learning Grounded Multilingual Multimodal Representations (D19-1)
Copied to clipboard
| Challenge: | Recent studies have advanced learning VSE under the monolingual setup. |
| Approach: | They propose a model with diverse multi-head attention to learn grounded multilingual multimodal representations by leveraging visual object detection. |
| Outcome: | The proposed model performs well in German-Image and English-Image matching tasks and in the Semantic Textual Similarity task with English descriptions of visual content. |
Similar Papers
A Visual Attention Grounding Neural Model for Multimodal Machine Translation (D18-1)
Copied to clipboard
| Challenge: | Existing approaches to multimodal machine translation do not integrate visual information into the translation process. |
| Approach: | They propose a multimodal machine translation model that utilizes parallel visual and textual information. |
| Outcome: | The proposed model outperforms existing methods on the Multi30K and Ambiguous COCO datasets. |
Aligning Multilingual Word Embeddings for Cross-Modal Retrieval Task (D19-64)
Copied to clipboard
| Challenge: | Existing methods to learn multimodal multilingual embeddings for text and image retrieval tasks are limited to English. |
| Approach: | They propose a new approach to learn multimodal multilingual embeddings for matching images and captions in two languages by combing two existing objective functions and adapting alignment between existing languages. |
| Outcome: | The proposed model achieves state-of-the-art in retrieval and caption-caption tasks while adapting existing language alignments. |
Aligning Multilingual Word Embeddings for Cross-Modal Retrieval Task (D19-66)
Copied to clipboard
| Challenge: | Existing methods to learn multimodal multilingual embeddings for text and image retrieval tasks are limited to English. |
| Approach: | They propose a new approach to learn multimodal multilingual embeddings for matching images and captions in two languages by combing two existing objective functions and adapting alignment between existing languages. |
| Outcome: | The proposed model achieves state-of-the-art in retrieval and caption-caption tasks while adapting existing language alignments. |
Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models (2021.naacl-main)
Copied to clipboard
| Challenge: | a new study examines zero-shot cross-lingual transfer of vision-language models . we study multilingual text-to-video search in non-English languages without annotations . |
| Approach: | They propose a Transformer-based model that learns contextual multilingual multimodal embeddings . they propose 'zero-shot cross-lingual transfer' to improve multilingual search . |
| Outcome: | The proposed model outperforms baselines on multilingual text-to-video search and multilingual image search on VTT and VATEX. |
Cross-lingual Cross-modal Pretraining for Multimodal Retrieval (2021.naacl-main)
Copied to clipboard
| Challenge: | Recent pretrained vision-language models have achieved impressive performance on cross-modal retrieval tasks in English. |
| Approach: | They propose a new approach to learn cross-lingual cross-modal representations for matching images and captions in multiple languages using an annotated corpus. |
| Outcome: | The proposed model achieves impressive performance on two multimodal multilingual image caption benchmarks: Multi30k with German captions and MSCOCO with Japanese captions. |
A Multi-task Approach to Learning Multilingual Representations (P18-2)
Copied to clipboard
| Challenge: | Using a multi-task model, we learn word and sentence embeddings in a single task. |
| Approach: | They propose a multi-task modeling approach that trains a skip-gram model and a cross-lingual sentence similarity model to learn word and sentence embeddings together. |
| Outcome: | The proposed model can learn word and sentence embeddings in a multilingual distributed representations of text using a cross-lingual sentence similarity model. |
Cross-lingual Visual Pre-training for Multimodal Machine Translation (2021.eacl-main)
Copied to clipboard
Ozan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Madhyastha, Erkut Erdem, Aykut Erdem, Lucia Specia
| Challenge: | Pre-trained language models have been shown to improve performance in many natural language tasks. |
| Approach: | They propose to combine cross-lingual and visual pre-training to learn visually-grounded cross-linguistic representations using masked region classification and three-way parallel vision & language corpora. |
| Outcome: | The proposed models obtain state-of-the-art performance when fine-tuned for multimodal machine translation. |
Assessing Multilingual Fairness in Pre-trained Multimodal Representations (2022.findings-acl)
Copied to clipboard
| Challenge: | Recent pre-trained multimodal models have shown exceptional capabilities towards connecting images and natural language. |
| Approach: | They propose two new fairness notions for pre-trained multimodal models that consider language as the fairness recipient. |
| Outcome: | The proposed models can be generalized to multilingualism by cross-lingual alignment . the results show that the models are individually fair across languages . |
Recognizing Multimodal Entailment (2021.acl-tutorials)
Copied to clipboard
Cesar Ilharco, Afsaneh Shirazi, Arjun Gopalan, Arsha Nagrani, Blaz Bratanic, Chris Bregler, Christina Funk, Felipe Ferreira, Gabriel Barcik, Gabriel Ilharco, Georg Osang, Jannis Bulian, Jared Frank, Lucas Smaira, Qin Cao, Ricardo Marino, Roma Patel, Thomas Leung, Vaiva Imbrasaite
| Challenge: | This tutorial introduces the multimodal entailment task for detecting semantic alignments . the task requires fine-grained understanding of visual and linguistic semantics questions . |
| Approach: | This tutorial introduces the multimodal entailment task to machine learning . it introduces a dataset for recognizing multimodal alignments . |
| Outcome: | This tutorial introduces the multimodal entailment task . it can be useful for detecting semantic alignments when a single modality alone is not enough . |
Grounding Multilingual Multimodal LLMs With Cultural Knowledge (2025.emnlp-main)
Copied to clipboard
| Challenge: | a new data-centric approach could address cultural gaps in multimodal large language models . despite being trained on billions of image-text pairs, today's models are biased towards English and Western data. |
| Approach: | They propose a data-centric approach that directly grounds MLLMs in cultural knowledge. |
| Outcome: | The proposed approach outperforms open-source models on cultural-focused benchmarks without degrading results on mainstream vision–language tasks. |