How to Adapt Pre-trained Vision-and-Language Models to a Text-only Input? (2022.coling-1)
Copied to clipboard
| Challenge: | Current language models have been criticised for learning language from text alone without connection between words and their meaning. |
| Approach: | They propose to train models on more sources than text to provide the lacking connection between words and their meanings. |
| Outcome: | The proposed model adaptation methods perform differently for different models and unimodal model counterparts perform on par with the VL models regardless of adaptation. |
Similar Papers
Does Vision-and-Language Pretraining Improve Lexical Grounding? (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Large pretrained language models (LMs) have been criticized for lack of grounding, i.e., connecting words to their meanings in the physical world. |
| Approach: | They compare vision-and-language (VL) models trained jointly on text and image or video data to find out how they compare to text-only counterparts. |
| Outcome: | The proposed model outperforms the text-only variants on a commonsense question answering task. |
Vision-Language Pretraining: Current Trends and the Future (2022.acl-tutorials)
Copied to clipboard
| Challenge: | Recent vision-language models are being used for downstream tasks that require large datasets and supervised datasets. |
| Approach: | They focus on recent vision-language pretraining paradigms and their strengths and shortcomings . they compare the different family of models used for vision- language pretraining . |
| Outcome: | This paper provides the background on image–language datasets, benchmarks, and modeling innovations before the multimodal pretraining area. |
Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models (2021.naacl-main)
Copied to clipboard
| Challenge: | a new study examines zero-shot cross-lingual transfer of vision-language models . we study multilingual text-to-video search in non-English languages without annotations . |
| Approach: | They propose a Transformer-based model that learns contextual multilingual multimodal embeddings . they propose 'zero-shot cross-lingual transfer' to improve multilingual search . |
| Outcome: | The proposed model outperforms baselines on multilingual text-to-video search and multilingual image search on VTT and VATEX. |
MAPL: Parameter-Efficient Adaptation of Unimodal Pre-Trained Models for Vision-Language Few-Shot Prompting (2023.eacl-main)
Copied to clipboard
| Challenge: | Large pre-trained models have proved to be remarkable zero- and (prompt-based) few-shot learners in unimodal vision and language tasks. |
| Approach: | They propose to use frozen unimodal models to learn a lightweight mapping between the representation spaces of unimod models using aligned image-text data. |
| Outcome: | The proposed method can generalize to unseen VL tasks from a few in-context examples while training orders of magnitude fewer parameters. |
Images in Language Space: Exploring the Suitability of Large Language Models for Vision & Language Tasks (2023.findings-acl)
Copied to clipboard
| Challenge: | Large language models have demonstrated robust performance on various language tasks using zero-shot or few-shot learning paradigms. |
| Approach: | They propose to use open-source, open-access language models to make visual input accessible to the model using separate verbalisation models. |
| Outcome: | The proposed model can handle visual input but also require strong reasoning component. |
Cross-lingual Visual Pre-training for Multimodal Machine Translation (2021.eacl-main)
Copied to clipboard
Ozan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Madhyastha, Erkut Erdem, Aykut Erdem, Lucia Specia
| Challenge: | Pre-trained language models have been shown to improve performance in many natural language tasks. |
| Approach: | They propose to combine cross-lingual and visual pre-training to learn visually-grounded cross-linguistic representations using masked region classification and three-way parallel vision & language corpora. |
| Outcome: | The proposed models obtain state-of-the-art performance when fine-tuned for multimodal machine translation. |
Match the Script, Adapt if Multilingual: Analyzing the Effect of Multilingual Pretraining on Cross-lingual Transferability (2022.acl-long)
Copied to clipboard
| Challenge: | Pretrained multilingual models enable zero-shot learning even for unseen languages . current multilingual model covers only a small subset of the world's languages - due to data sparsity, they are not likely to obtain good results for many lowresource languages. |
| Approach: | They ask: how does the number of pretraining languages influence zero-shot learning for unseen languages? do the findings change if the languages used for pretraining are all related? |
| Outcome: | The results show that pretrained models can zero-shot learn for unseen languages even for limited amounts even for low-resource languages. |
Can Monolingual Pretrained Models Help Cross-Lingual Classification? (2020.aacl-main)
Copied to clipboard
| Challenge: | Multilingual pretrained language models have shown impressive results for cross-lingual transfer, but due to the constant model capacity, multilingual pre-training usually lags behind the monolingual competitors. |
| Approach: | They propose to transfer the knowledge from monolingual pretrained models to multilingual ones to improve zero-shot cross-lingual classification by using machine translation systems. |
| Outcome: | The proposed methods outperform vanilla multilingual fine-tuning on two cross-lingual classification benchmarks. |
Unifying Cross-Lingual and Cross-Modal Modeling Towards Weakly Supervised Multilingual Vision-Language Pre-training (2023.acl-long)
Copied to clipboard
| Challenge: | Existing studies address the problem of translating English data into other languages, but they are limited in form and scale. |
| Approach: | They propose a framework to unify cross-lingual and cross-modal pre-training by using English data. |
| Outcome: | The proposed framework unifies cross-lingual and cross-modal pre-training on different data. |
Stop Pre-Training: Adapt Visual-Language Models to Unseen Languages (2023.acl-short)
Copied to clipboard
| Challenge: | Existing studies have shown that the pre-training in English does not transfer well to other languages in a zero-shot setting. |
| Approach: | They propose a simple yet efficient approach to adapt VLP to unseen languages using MPLM. |
| Outcome: | The proposed approach outperforms state-of-the-art models without large parallel corpora across three tasks. |