Challenge: Large-scale "natural language supervision" using image captions has enabled the first "zero-shot" AI image classifiers, which allow users to create their own image classes using natural language, yet outperform supervised models on common language-and-image tasks.
Approach: They compare the geometry and semantic properties of contextualized English language representations formed by GPT-2 and CLIP, a zero-shot multimodal image classifier which adapts the GPT2 architecture to encode image captions.
Outcome: The proposed classifier outperforms GPT-2 on word-level semantic intrinsic evaluation tasks and achieves a new corpus-based state of the art for the RG65 evaluation.

Similar Papers

On the Sentence Embeddings from Pre-trained Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Pre-trained contextual representations like BERT have been widely used for NLP tasks.
Approach: They propose to transform anisotropic sentence embedding distribution to smooth and isotropic Gaussian distribution by normalizing flows that are learned with an unsupervised objective.
Outcome: The proposed method achieves significant performance gains over state-of-the-art embeddings on a variety of semantic textual similarity tasks.
Learning Visually-Grounded Semantics from Contrastive Adversarial Samples (C18-1)

Copied to clipboard

Challenge: Existing frameworks for grounding distributional representations of texts on the visual domain are limited . effective and efficient grounding of distributional embeddings remains challenging .
Approach: They propose to ground distributional representations of texts on the visual domain using visual-semantic embeddings.
Outcome: The proposed model improves on a diverse set of downstream tasks and defends known-type adversarial attacks.
Delving into the Openness of CLIP (2023.findings-acl)

Copied to clipboard

Challenge: Contrastive Language-Image Pre-training (CLIP) allows for open-vocabulary visual recognition, where the model can recognize images from an open class set in a zero-shot manner.
Approach: They propose to use image classification as an image-to-text matching task instead of discrete category IDs to achieve open-vocabulary visual recognition.
Outcome: The proposed model can recognize images from an open vocabulary in a zero-shot manner, but its performance deteriorates as the vocabulary expands.
Injecting Wiktionary to improve token-level contextual representations using contrastive learning (2024.eacl-short)

Copied to clipboard

Challenge: lexical semantics tasks require contextual word embeddings that are not blind to context, despite the fact that vectors of the same meaning are too different.
Approach: They propose to fine-tune pre-trained language models by using automatically self-augmented examples to target contextual word embeddings.
Outcome: The proposed method achieves significant improvements on the original WiC test set and in two new tests.
Sub-Sentence Encoder: Contrastive Learning of Propositional Semantic Representations (2024.naacl-long)

Copied to clipboard

Challenge: Sentence embeddings are typically learned to recognize the semantic relation between two text inputs.
Approach: They introduce a contrastively-learned contextual embedding model for fine-grained semantic representation of text.
Outcome: The proposed model is able to produce contextual embeddings corresponding to different atomic propositions, i.e. semantic equivalence between propositions across different text sequences.
Lexicon-Level Contrastive Visual-Grounding Improves Language Modeling (2024.findings-acl)

Copied to clipboard

Challenge: Neural language models (LMs) are trained on orders of magnitude more language data than human language learners receive, but without supervision from other sensory modalities that play a crucial role in human learning.
Approach: They propose a grounded language learning procedure that leverages visual supervision to improve textual representations.
Outcome: The proposed procedure outperforms standard language-only models in terms of learning efficiency in small and developmentally plausible data regimes and improves perplexity by around 5% on multiple language modeling tasks compared to other models trained on the same amount of text data.
On the Language Encoder of Contrastive Cross-modal Models (2024.findings-acl)

Copied to clipboard

Challenge: Pretrained audio-language models such as AudioCLIP and AudioCLAP have shown promising results on vision-language (VL) tasks.
Approach: They extensively evaluate how unsupervised and supervised sentence embedding training affect language encoder quality and cross-modal task performance.
Outcome: The proposed model improves on visual-language (VL) and audio-language tasks when the amount of training data is large.
ComCLIP: Training-Free Compositional Image and Text Matching (2024.naacl-long)

Copied to clipboard

Challenge: erroneous semantics of individual entities are essentially confounders that cause the matching failure.
Approach: They propose a training-free compositional CLIP model which disentangles input images into subjects, objects, and action subimages and composes CLIP’s vision encoder and text encoder to perform evolving matching over compositional text embedding and subimage embeddments.
Outcome: The proposed model mitigates spurious correlations introduced by the pretrained CLIP models and dynamically evaluates the importance of each component.
VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding (2021.emnlp-main)

Copied to clipboard

Challenge: Recent work adopts a "pre-training + fine-tuning" approach for zero-shot transfer to end tasks without fine- tuning.
Approach: They propose a contrastive approach to pre-train a transformer model for zero-shot video and text understanding without using any labels on downstream tasks.
Outcome: The proposed model outperforms supervised approaches on downstream tasks and outperformed previous approaches.
Infusing Finetuning with Semantic Dependencies (2021.tacl-1)

Copied to clipboard

Challenge: Several diagnostics help to localize the benefits of our approach.
Approach: They apply convolutional graph encoders to integrate semantic parses into task-specific finetuning.
Outcome: The proposed approach yields benefits to natural language understanding (NLU) tasks in the GLUE benchmark.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations