| Challenge: | Existing multimodal corpora lack the ability to be used in multilingual or non-English scenarios. |
| Approach: | They extend a Flickr30k Entities image-caption dataset with Japanese translations to provide a multilingual corpus. |
| Outcome: | The proposed dataset is the first multilingual image-caption dataset with Japanese translations. |
Similar Papers
Adding Syntactic Annotations to Flickr30k Entities Corpus for Multimodal Ambiguous Prepositional-Phrase Attachment Resolution (L18-1)
Copied to clipboard
| Challenge: | Using visual features extracted from an image, we propose to study the joint processing of image and language features for the Preposition-Phrase attachment disambiguation task. |
| Approach: | They propose to add syntactic annotations to the captions of the Flickr30k Entities corpus to study the joint processing of image and language features for the Preposition-Phrase attachment disambiguation task. |
| Outcome: | The proposed framework is based on the captions of the Flickr30k Entities corpus and is automatically projected on their French and German translations. |
Cross-lingual Cross-modal Pretraining for Multimodal Retrieval (2021.naacl-main)
Copied to clipboard
| Challenge: | Recent pretrained vision-language models have achieved impressive performance on cross-modal retrieval tasks in English. |
| Approach: | They propose a new approach to learn cross-lingual cross-modal representations for matching images and captions in multiple languages using an annotated corpus. |
| Outcome: | The proposed model achieves impressive performance on two multimodal multilingual image caption benchmarks: Multi30k with German captions and MSCOCO with Japanese captions. |
Framed Multi30K: A Frame-Based Multimodal-Multilingual Dataset (2024.lrec-main)
Copied to clipboard
Marcelo Viridiano, Arthur Lorenzi, Tiago Timponi Torrent, Ely E. Matos, Adriana S. Pagano, Natália Sathler Sigiliano, Maucha Gamonal, Helen de Andrade Abreu, Lívia Vicente Dutra, Mairon Samagaio, Mariane Carvalho, Franciany Campos, Gabrielly Azalim, Bruna Mazzei, Mateus Fonseca de Oliveira, Ana Carolina Luz, Livia Padua Ruiz, Júlia Bellei, Amanda Pestana, Josiane Costa, Iasmin Rabelo, Anna Beatriz Silva, Raquel Roza, Mariana Souza Mota, Igor Oliveira, Márcio Henrique Pelegrino de Freitas
| Challenge: | Recent advances in image-captioning datasets combine image and language to solve a diverse range of tasks. |
| Approach: | They propose a Brazilian Portuguese multimodal-multilingual dataset that extends the Multi30K dataset with 158,915 original Brazilian Portuguese descriptions and 30,104 Brazilian Portuguese translations. |
| Outcome: | The proposed dataset adds 2,677,613 frame evocation labels to the 158,915 English descriptions and to the ones created for Brazilian Portuguese. |
Multilingual Image Corpus – Towards a Multimodal and Multilingual Dataset (2022.lrec-1)
Copied to clipboard
| Challenge: | The goal of the project Multilingual Image Corpus is to provide a large image dataset with annotated objects and object descriptions in 24 languages. |
| Approach: | They propose to provide a large image dataset with annotated objects and object descriptions in 24 languages. |
| Outcome: | The project provides a large image dataset with annotated objects and object descriptions in 24 languages. |
Learning Translations via Images with a Massively Multilingual Image Dataset (P18-1)
Copied to clipboard
| Challenge: | Existing datasets for learning translations of words are limited to a few high-resource languages and unrealistically easy settings. |
| Approach: | They propose a large-scale multilingual corpus of images labeled with the word they represent to facilitate translation research. |
| Outcome: | The proposed method improves on an unsupervised technique that has been limited to a few languages and unrealistic settings. |
Entity Linking in 100 Languages (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to multilingual entity linking are cross-lingual, with a focus on zero-shot evaluation. |
| Approach: | They propose a new formulation for multilingual entity linking where language-specific mentions resolve to a language-agnostic Knowledge Base. |
| Outcome: | The proposed model outperforms state-of-the-art models on a large multilingual dataset and shows that frequency-based analysis provided key insights for the model and training enhancements. |
Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on multilingual image captioning have been hampered by a lack of high-quality evaluation datasets. |
| Approach: | They present a dataset of 3600 images annotated with human-generated captions in 36 languages. |
| Outcome: | The proposed dataset shows that it is feasible to build multilingual image captioning models trained on machine-translated data. |
Developing Japanese CLIP Models Leveraging an Open-weight LLM for Large-scale Dataset Translation (2025.naacl-srw)
Copied to clipboard
| Challenge: | lack of large-scale open Japanese image-text pairs poses a significant barrier to the development of vision-language models. |
| Approach: | They construct large-scale Japanese image-text pairs using machine translation and pre-trained CLIP models on a Japanese dataset. |
| Outcome: | The results show that pre-trained models achieve competitive average scores on Japanese culture tasks compared to models of similar size. |
Aligning Multilingual Word Embeddings for Cross-Modal Retrieval Task (D19-66)
Copied to clipboard
| Challenge: | Existing methods to learn multimodal multilingual embeddings for text and image retrieval tasks are limited to English. |
| Approach: | They propose a new approach to learn multimodal multilingual embeddings for matching images and captions in two languages by combing two existing objective functions and adapting alignment between existing languages. |
| Outcome: | The proposed model achieves state-of-the-art in retrieval and caption-caption tasks while adapting existing language alignments. |
Aligning Multilingual Word Embeddings for Cross-Modal Retrieval Task (D19-64)
Copied to clipboard
| Challenge: | Existing methods to learn multimodal multilingual embeddings for text and image retrieval tasks are limited to English. |
| Approach: | They propose a new approach to learn multimodal multilingual embeddings for matching images and captions in two languages by combing two existing objective functions and adapting alignment between existing languages. |
| Outcome: | The proposed model achieves state-of-the-art in retrieval and caption-caption tasks while adapting existing language alignments. |