Challenge: Existing frameworks for grounding distributional representations of texts on the visual domain are limited . effective and efficient grounding of distributional embeddings remains challenging .
Approach: They propose to ground distributional representations of texts on the visual domain using visual-semantic embeddings.
Outcome: The proposed model improves on a diverse set of downstream tasks and defends known-type adversarial attacks.

Similar Papers

Lexicon-Level Contrastive Visual-Grounding Improves Language Modeling (2024.findings-acl)

Copied to clipboard

Challenge: Neural language models (LMs) are trained on orders of magnitude more language data than human language learners receive, but without supervision from other sensory modalities that play a crucial role in human learning.
Approach: They propose a grounded language learning procedure that leverages visual supervision to improve textual representations.
Outcome: The proposed procedure outperforms standard language-only models in terms of learning efficiency in small and developmentally plausible data regimes and improves perplexity by around 5% on multiple language modeling tasks compared to other models trained on the same amount of text data.
Attacking Visual Language Grounding with Adversarial Examples: A Case Study on Neural Image Captioning (P18-1)

Copied to clipboard

Challenge: Visual language grounding is widely studied in modern neural image captioning systems . a novel algorithm for crafting adversarial examples in image captions is proposed .
Approach: They propose an algorithm to craft adversarial examples in machine vision and perception . their approach provides two evaluation approaches to check if they can mislead systems .
Outcome: The proposed algorithm can craft visually-similar adversarial examples with randomly targeted captions or keywords, and the results are transferable to other image captioning systems.
Improving Visual-Semantic Embedding with Adaptive Pooling and Optimization Objective (2023.eacl-main)

Copied to clipboard

Challenge: Recent VSE models combine simple pooling methods with hard triplet loss to improve performance.
Approach: They propose an adaptive pooling strategy that allows the model to learn how to aggregate features through a combination of simple pooling methods.
Outcome: The proposed strategy outperforms current state-of-the-art systems on image-to-text and text-toimage retrieval.
Contrastive Visual Semantic Pretraining Magnifies the Semantics of Natural Language Representations (2022.acl-long)

Copied to clipboard

Challenge: Large-scale "natural language supervision" using image captions has enabled the first "zero-shot" AI image classifiers, which allow users to create their own image classes using natural language, yet outperform supervised models on common language-and-image tasks.
Approach: They compare the geometry and semantic properties of contextualized English language representations formed by GPT-2 and CLIP, a zero-shot multimodal image classifier which adapts the GPT2 architecture to encode image captions.
Outcome: The proposed classifier outperforms GPT-2 on word-level semantic intrinsic evaluation tasks and achieves a new corpus-based state of the art for the RG65 evaluation.
Learning Visually Grounded Sentence Representations (N18-1)

Copied to clipboard

Challenge: Unsupervised sentence representation models suffer from the grounding problem because of lack of association between symbols and external information.
Approach: They train a sentence encoder to predict image features of a caption and use them as sentence representations.
Outcome: The proposed model improves on word embeddings and word representations on standard benchmarks.
Measuring Social Biases in Grounded Vision and Language Embeddings (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to measure social biases in word embeddings are limited to visually grounded word embeds . a new study generalizes word embedment associations to visually ground word embeddas .
Approach: They generalize word embeddings' biases to visually grounded word embeds . they propose two generalizations that answer questions about how biase, language, and vision interact .
Outcome: The proposed measures are applied to a new dataset that includes 10,228 images from COCO, Conceptual Captions, and Google Images.
SimCSE: Simple Contrastive Learning of Sentence Embeddings (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for learning universal sentence embeddings are based on unsupervised approaches with only dropout as noise.
Approach: They propose an unsupervised approach that takes an input sentence and predicts itself in a contrastive objective with only standard dropout used as noise.
Outcome: The proposed framework performs on par with previous supervised approaches and can produce superior sentence embeddings from unlabeled or labeled data.
Illustrative Language Understanding: Large-Scale Visual Grounding with Image Search (P18-1)

Copied to clipboard

Challenge: a large-scale lookup operation to ground language via ‘snapshots’ of our physical world accessed through image search is currently used to learn word representations.
Approach: They propose a large-scale lookup operation to ground language via ‘snapshots’ of our physical world accessed through image search.
Outcome: The proposed model is based on a large-scale lookup operation to ground language using image search.
Adversarial Contrastive Estimation (P18-1)

Copied to clipboard

Challenge: Noise contrastive estimation (NCE) is a general strategy used in word embeddings and translations for knowledge graphs.
Approach: They propose to augment negative sampler into mixture distribution with adversarially learned sampler and to combine it with noise contrastive estimation (NCE) they observe faster convergence and improved results on multiple metrics.
Outcome: The proposed model performs better on word embeddings, order embedds and knowledge graph embeddments and faster convergence and improved results on multiple metrics.
Visually Grounded Compound PCFGs (2020.emnlp-main)

Copied to clipboard

Challenge: Existing work on visual groundings for language understanding has been drawing much attention.
Approach: They propose to use an extension of probabilistic context-free grammar model to do fully-differentiable end-to-end visually grounded learning.
Outcome: The proposed model outperforms the previous grounded model and significantly outperformed the previous model on the MSCOCO test captions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations