Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations (2020.emnlp-main)
Copied to clipboard
| Challenge: | A major challenge in visually grounded language generation is to build robust benchmark datasets and models that can generalize well in real-world settings. |
| Approach: | They propose to use visual attention to build robust benchmark datasets and models that can generalize well in real-world settings. |
| Outcome: | The proposed models show that human-generated references vary drastically in different datasets/tasks, revealing the nature of each task. |
Similar Papers
Cross-Lingual and Cross-Cultural Variation in Image Descriptions (2025.naacl-long)
Copied to clipboard
| Challenge: | Behavioural and cognitive studies report cultural effects on perception, but these are limited in scope and hard to replicate. |
| Approach: | They develop a method to accurately identify entities mentioned in captions and present in images, then measure how they vary across languages. |
| Outcome: | The proposed method corroborates previous studies showing that languages that are geographically or genetically closer mention entities more frequently than others. |
Measuring Social Biases in Grounded Vision and Language Embeddings (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing methods to measure social biases in word embeddings are limited to visually grounded word embeds . a new study generalizes word embedment associations to visually ground word embeddas . |
| Approach: | They generalize word embeddings' biases to visually grounded word embeds . they propose two generalizations that answer questions about how biase, language, and vision interact . |
| Outcome: | The proposed measures are applied to a new dataset that includes 10,228 images from COCO, Conceptual Captions, and Google Images. |
Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Vision Language Models (LVLMs) have shown impressive performance on various vision-language tasks. |
| Approach: | They propose a benchmark framework for evaluating Visual Variation Robustness of Large Vision Language Models that incorporates automated evaluation dataset generation and principled metrics for thorough robustness assessment. |
| Outcome: | The proposed framework identifies a vulnerability to visual variations affecting even advanced models that excel at complex vision-language tasks but significantly underperform on simple tasks like object recognition. |
Describing Images Fast and Slow: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing models of visuo-linguistic variation are weak to moderately trained to capture such a variation in visual outputs. |
| Approach: | They use a corpus of Dutch image descriptions with eye-tracking data to investigate the nature of the variation in visuo-linguistic signals. |
| Outcome: | The proposed model lacks biases about what makes a stimulus complex for humans and what leads to variations in human outputs. |
Curious Case of Language Generation Evaluation Metrics: A Cautionary Tale (2020.coling-main)
Copied to clipboard
| Challenge: | a few popular metrics are still used to evaluate language generation systems despite their known limitations. |
| Approach: | They propose to use automatic metrics to evaluate language generation systems . they show that they prefer system outputs to human-authored texts . |
| Outcome: | The proposed metrics are insensitive to correct translations of rare words and can yield high scores when given a single sentence as system output for the entire test set. |
Learning Visually-Grounded Semantics from Contrastive Adversarial Samples (C18-1)
Copied to clipboard
| Challenge: | Existing frameworks for grounding distributional representations of texts on the visual domain are limited . effective and efficient grounding of distributional embeddings remains challenging . |
| Approach: | They propose to ground distributional representations of texts on the visual domain using visual-semantic embeddings. |
| Outcome: | The proposed model improves on a diverse set of downstream tasks and defends known-type adversarial attacks. |
Visual Referring Expression Recognition: What Do Systems Actually Learn? (N18-2)
Copied to clipboard
| Challenge: | Existing systems for referring expression recognition ignore linguistic structure, instead relying on shallow correlations introduced by unintended biases in the data selection and annotation process. |
| Approach: | They propose to use a system trained on the input image without the input referring expression to achieve a precision of 71.2% in top-2 predictions. |
| Outcome: | The proposed model can achieve 71.2% accuracy on the input image without the input referring expression and 84.2% on the object category given the input. |
What Are We Measuring When We Evaluate Large Vision-Language Models? An Analysis of Latent Factors and Biases (2024.naacl-long)
Copied to clipboard
| Challenge: | Vision-language models have broad competence that is difficult to evaluate . current evaluation benchmarks focus on only assessing one or a few capabilities . |
| Approach: | They perform a large-scale transfer learning experiment to discover latent VL skills from data. |
| Outcome: | The results suggest that factor analysis can identify reasonable yet surprising VL skill factors . the results contribute to the design of balanced and broad-coverage vision-language evaluation methods. |
Learning Visually Grounded Sentence Representations (N18-1)
Copied to clipboard
| Challenge: | Unsupervised sentence representation models suffer from the grounding problem because of lack of association between symbols and external information. |
| Approach: | They train a sentence encoder to predict image features of a caption and use them as sentence representations. |
| Outcome: | The proposed model improves on word embeddings and word representations on standard benchmarks. |
Towards Explainable Evaluation of Language Models on the Semantic Similarity of Visual Concepts (2022.coling-1)
Copied to clipboard
Maria Lymperaiou, George Manoliadis, Orfeas Menis Mastromichalakis, Edmund G. Dervakos, Giorgos Stamou
| Challenge: | Recent advances in NLP research have focused on robustness and explainability issues of their evaluation strategies. |
| Approach: | They propose to use pre-trained transformers to evaluate semantic similarity for visual vocabularies . they propose to provide explainable metrics for understanding the quality of retrieved instances . |
| Outcome: | The proposed metrics highlight inabilities of widely used evaluation methods and highlight weaknesses in learned linguistic representations. |