Challenge: Large pretrained language models (LMs) have been criticized for lack of grounding, i.e., connecting words to their meanings in the physical world.
Approach: They compare vision-and-language (VL) models trained jointly on text and image or video data to find out how they compare to text-only counterparts.
Outcome: The proposed model outperforms the text-only variants on a commonsense question answering task.

Similar Papers

How to Adapt Pre-trained Vision-and-Language Models to a Text-only Input? (2022.coling-1)

Copied to clipboard

Challenge: Current language models have been criticised for learning language from text alone without connection between words and their meaning.
Approach: They propose to train models on more sources than text to provide the lacking connection between words and their meanings.
Outcome: The proposed model adaptation methods perform differently for different models and unimodal model counterparts perform on par with the VL models regardless of adaptation.
Are Visual-Linguistic Models Commonsense Knowledge Bases? (2022.coling-1)

Copied to clipboard

Challenge: PTLMs are used to extract knowledge from text on demand.
Approach: They compare visual-linguistic and language-only visual-language models in a zero-shot commonsense question answering inference task.
Outcome: The proposed models are highly promising on certain types of commonsense knowledge associated with the visual world.
Vision-Language Pretraining: Current Trends and the Future (2022.acl-tutorials)

Copied to clipboard

Challenge: Recent vision-language models are being used for downstream tasks that require large datasets and supervised datasets.
Approach: They focus on recent vision-language pretraining paradigms and their strengths and shortcomings . they compare the different family of models used for vision- language pretraining .
Outcome: This paper provides the background on image–language datasets, benchmarks, and modeling innovations before the multimodal pretraining area.
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models (2024.eacl-long)

Copied to clipboard

Challenge: Existing vision-and-language models perform better on multimodal tasks, but there is little understanding of how multimodal learning can help visual representations.
Approach: They conduct a probing analysis of visual representations in existing vision-and-language models and vision-only models by probing on a broad range of tasks.
Outcome: The proposed model improves vision-and-language models on label and attribute prediction tasks while vision-only models are stronger on dense prediction tasks.
Visual Grounding Helps Learn Word Meanings in Low-Data Regimes (2024.naacl-long)

Copied to clipboard

Challenge: Modern neural language models (LMs) require distinctly un-human-like ways to achieve these results.
Approach: They train a diverse set of LM architectures with and without auxiliary visual supervision on datasets of varying scales.
Outcome: The proposed models exhibit better learning of syntactic categories, lexical relations, semantic features, word similarity and alignment with human neural representations.
Cross-lingual Visual Pre-training for Multimodal Machine Translation (2021.eacl-main)

Copied to clipboard

Challenge: Pre-trained language models have been shown to improve performance in many natural language tasks.
Approach: They propose to combine cross-lingual and visual pre-training to learn visually-grounded cross-linguistic representations using masked region classification and three-way parallel vision & language corpora.
Outcome: The proposed models obtain state-of-the-art performance when fine-tuned for multimodal machine translation.
Does Vision Accelerate Hierarchical Generalization in Neural Language Learners? (2025.coling-main)

Copied to clipboard

Challenge: Neural language models (LMs) are arguably less data-efficient than humans from a language acquisition perspective.
Approach: They investigate the advantage of grounded language acquisition over visual input to improve syntactic generalization.
Outcome: The proposed model is less efficient than humans in language acquisition . it shows that visual input helps syntactic generalization, but not vision .
Grounded Concreteness: Human-Like Concreteness Sensitivity in Vision–Language Models (2026.findings-acl)

Copied to clipboard

Challenge: a long tradition in cognitive science treats concreteness as a graded dimension of conceptual representation . concrete words benefit from richer sensory codes and exhibit robust behavioral advantages over abstract words .
Approach: They compare vision-language models with text-only large language models to test their concreteness . they find that VLMs show more human-like sensitivity to concreteness than LLMs .
Outcome: The proposed model-based training improves on the Llama text backbones and Llma Vision counterparts.
Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models? (2025.findings-emnlp)

Copied to clipboard

Challenge: Despite significant progress in multimodal language models, it remains unclear whether visual grounding enhances their understanding of embodied knowledge compared to text-only models.
Approach: They propose to assess vision-language models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions.
Outcome: The proposed benchmark assesses the models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions.
Multimodal Pretraining Unmasked: A Meta-Analysis and a Unified Framework of Vision-and-Language BERTs (2021.tacl-1)

Copied to clipboard

Challenge: Large-scale pretraining and task-specific fine-tuning are now the standard methodology for many tasks in computer vision and natural language processing.
Approach: They propose to combine two types of vision and language BERTs to create a theoretical framework that can be unified under different theoretical frameworks.
Outcome: The proposed models can be classified into single-stream or dual-stream encoders and are unified under a single theoretical framework.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations