Challenge: Embedding a sentence into a point in a highdimensional continuous space plays a foundational role in the natural language processing.
Approach: They propose to use contrastive loss to fine-tune sentences by inverse word frequency . they also show that more informative words receive greater weight than less informative ones .
Outcome: The proposed method improves the performance of sentence embeddings by weighing them based on information-theoretic quantities.

Similar Papers

Sub-Sentence Encoder: Contrastive Learning of Propositional Semantic Representations (2024.naacl-long)

Copied to clipboard

Challenge: Sentence embeddings are typically learned to recognize the semantic relation between two text inputs.
Approach: They introduce a contrastively-learned contextual embedding model for fine-grained semantic representation of text.
Outcome: The proposed model is able to produce contextual embeddings corresponding to different atomic propositions, i.e. semantic equivalence between propositions across different text sequences.
On the Language Encoder of Contrastive Cross-modal Models (2024.findings-acl)

Copied to clipboard

Challenge: Pretrained audio-language models such as AudioCLIP and AudioCLAP have shown promising results on vision-language (VL) tasks.
Approach: They extensively evaluate how unsupervised and supervised sentence embedding training affect language encoder quality and cross-modal task performance.
Outcome: The proposed model improves on visual-language (VL) and audio-language tasks when the amount of training data is large.
Improving Contrastive Learning of Sentence Embeddings with Focal InfoNCE (2023.findings-emnlp)

Copied to clipboard

Challenge: SimCSE does not fully exploit the potential of hard negative samples in contrastive learning.
Approach: They propose an unsupervised contrastive learning framework that combines SimCSE with hard negative mining to enhance the quality of sentence embeddings.
Outcome: The proposed framework improves sentence embeddings on various STS benchmarks in terms of Spearman’s correlation, representation alignment and uniformity.
Composition-contrastive Learning for Sentence Embeddings (2023.acl-long)

Copied to clipboard

Challenge: Recent work shows potential to learn vector representations from unlabelled data without task-specific fine-tuning.
Approach: They propose to maximize alignment between textual embeddings and a composition of their phrasal constituents.
Outcome: The proposed approach improves on similarity tasks comparable to state-of-the-art approaches.
Batch-Softmax Contrastive Loss for Pairwise Sentence Scoring Tasks (2022.naacl-main)

Copied to clipboard

Challenge: Recent advances in machine learning have led to the use of contrastive loss for representation learning.
Approach: They propose to use batch-softmax contrastive loss to train pairwise sentence embeddings . they propose to take a batch-softermax contrastitive loss and train it with different loss functions .
Outcome: The proposed model improves on a number of datasets and pairwise sentence scoring tasks.
Differentiable Data Augmentation for Contrastive Sentence Representation Learning (2022.emnlp-main)

Copied to clipboard

Challenge: a contrastive learning framework is used to fine-tune pre-trained language models with unlabeled sentences or labeled sentences.
Approach: They propose a method that makes hard positives from unlabeled sentences . they use a prefix attached to a model to allow for differentiable data augmentation .
Outcome: The proposed method yields significant improvements over existing methods under semi-supervised and supervised settings.
Learning Visually Grounded Sentence Representations (N18-1)

Copied to clipboard

Challenge: Unsupervised sentence representation models suffer from the grounding problem because of lack of association between symbols and external information.
Approach: They train a sentence encoder to predict image features of a caption and use them as sentence representations.
Outcome: The proposed model improves on word embeddings and word representations on standard benchmarks.
InfoCSE: Information-aggregated Contrastive Learning of Sentence Embeddings (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on contrastive learning for sentence embeddings are weak . researchers have started to use contrastive training to learn better unsupervised sentences.
Approach: They propose an information-aggregated contrastive learning framework for learning unsupervised sentence embeddings.
Outcome: The proposed framework outperforms SimCSE on several benchmark datasets w.r.t the semantic text similarity task.
Smoothed Contrastive Learning for Unsupervised Sentence Embedding (2022.coling-1)

Copied to clipboard

Challenge: Unsupervised contrastive sentence embedding models use InfoNCE loss function . increasing batch size leads to performance degradation when it exceeds threshold .
Approach: They propose a simple smoothing strategy upon the InfoNCE loss function to reduce the number of false-negative pairs in a batch without increasing the batch size.
Outcome: The proposed smoothing strategy improves unsupervised SimCSE on semantic similarity tasks.
SimCSE: Simple Contrastive Learning of Sentence Embeddings (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for learning universal sentence embeddings are based on unsupervised approaches with only dropout as noise.
Approach: They propose an unsupervised approach that takes an input sentence and predicts itself in a contrastive objective with only standard dropout used as noise.
Outcome: The proposed framework performs on par with previous supervised approaches and can produce superior sentence embeddings from unlabeled or labeled data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations