Challenge: A major challenge in visually grounded language generation is to build robust benchmark datasets and models that can generalize well in real-world settings.
Approach: They propose to use visual attention to build robust benchmark datasets and models that can generalize well in real-world settings.
Outcome: The proposed models show that human-generated references vary drastically in different datasets/tasks, revealing the nature of each task.

Similar Papers

Cross-Lingual and Cross-Cultural Variation in Image Descriptions (2025.naacl-long)

Copied to clipboard

Challenge: Behavioural and cognitive studies report cultural effects on perception, but these are limited in scope and hard to replicate.
Approach: They develop a method to accurately identify entities mentioned in captions and present in images, then measure how they vary across languages.
Outcome: The proposed method corroborates previous studies showing that languages that are geographically or genetically closer mention entities more frequently than others.
Measuring Social Biases in Grounded Vision and Language Embeddings (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to measure social biases in word embeddings are limited to visually grounded word embeds . a new study generalizes word embedment associations to visually ground word embeddas .
Approach: They generalize word embeddings' biases to visually grounded word embeds . they propose two generalizations that answer questions about how biase, language, and vision interact .
Outcome: The proposed measures are applied to a new dataset that includes 10,228 images from COCO, Conceptual Captions, and Google Images.
Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward (2025.findings-acl)

Copied to clipboard

Challenge: Large Vision Language Models (LVLMs) have shown impressive performance on various vision-language tasks.
Approach: They propose a benchmark framework for evaluating Visual Variation Robustness of Large Vision Language Models that incorporates automated evaluation dataset generation and principled metrics for thorough robustness assessment.
Outcome: The proposed framework identifies a vulnerability to visual variations affecting even advanced models that excel at complex vision-language tasks but significantly underperform on simple tasks like object recognition.
Describing Images Fast and Slow: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes (2024.eacl-long)

Copied to clipboard

Challenge: Existing models of visuo-linguistic variation are weak to moderately trained to capture such a variation in visual outputs.
Approach: They use a corpus of Dutch image descriptions with eye-tracking data to investigate the nature of the variation in visuo-linguistic signals.
Outcome: The proposed model lacks biases about what makes a stimulus complex for humans and what leads to variations in human outputs.
Curious Case of Language Generation Evaluation Metrics: A Cautionary Tale (2020.coling-main)

Copied to clipboard

Challenge: a few popular metrics are still used to evaluate language generation systems despite their known limitations.
Approach: They propose to use automatic metrics to evaluate language generation systems . they show that they prefer system outputs to human-authored texts .
Outcome: The proposed metrics are insensitive to correct translations of rare words and can yield high scores when given a single sentence as system output for the entire test set.
Learning Visually-Grounded Semantics from Contrastive Adversarial Samples (C18-1)

Copied to clipboard

Challenge: Existing frameworks for grounding distributional representations of texts on the visual domain are limited . effective and efficient grounding of distributional embeddings remains challenging .
Approach: They propose to ground distributional representations of texts on the visual domain using visual-semantic embeddings.
Outcome: The proposed model improves on a diverse set of downstream tasks and defends known-type adversarial attacks.
Visual Referring Expression Recognition: What Do Systems Actually Learn? (N18-2)

Copied to clipboard

Challenge: Existing systems for referring expression recognition ignore linguistic structure, instead relying on shallow correlations introduced by unintended biases in the data selection and annotation process.
Approach: They propose to use a system trained on the input image without the input referring expression to achieve a precision of 71.2% in top-2 predictions.
Outcome: The proposed model can achieve 71.2% accuracy on the input image without the input referring expression and 84.2% on the object category given the input.
What Are We Measuring When We Evaluate Large Vision-Language Models? An Analysis of Latent Factors and Biases (2024.naacl-long)

Copied to clipboard

Challenge: Vision-language models have broad competence that is difficult to evaluate . current evaluation benchmarks focus on only assessing one or a few capabilities .
Approach: They perform a large-scale transfer learning experiment to discover latent VL skills from data.
Outcome: The results suggest that factor analysis can identify reasonable yet surprising VL skill factors . the results contribute to the design of balanced and broad-coverage vision-language evaluation methods.
Learning Visually Grounded Sentence Representations (N18-1)

Copied to clipboard

Challenge: Unsupervised sentence representation models suffer from the grounding problem because of lack of association between symbols and external information.
Approach: They train a sentence encoder to predict image features of a caption and use them as sentence representations.
Outcome: The proposed model improves on word embeddings and word representations on standard benchmarks.
Towards Explainable Evaluation of Language Models on the Semantic Similarity of Visual Concepts (2022.coling-1)

Copied to clipboard

Challenge: Recent advances in NLP research have focused on robustness and explainability issues of their evaluation strategies.
Approach: They propose to use pre-trained transformers to evaluate semantic similarity for visual vocabularies . they propose to provide explainable metrics for understanding the quality of retrieved instances .
Outcome: The proposed metrics highlight inabilities of widely used evaluation methods and highlight weaknesses in learned linguistic representations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations