Challenge: Existing word embedding methods for natural language processing are limited in their ability to produce dense word embeds.
Approach: They propose a word embedding SentiVec which is infused with sentiment information from a lexical resource and outperforms baselines on subjectivity-sensitive tasks.
Outcome: The proposed word embedding SentiVec outperforms baselines on subjectivity-sensitive tasks.

Similar Papers

Analyzing the Surprising Variability in Word Embedding Stability Across Languages (2021.emnlp-main)

Copied to clipboard

Challenge: Word embeddings are powerful representations that form the foundation of many natural language processing architectures.
Approach: They explore word embedding stability in a wide range of languages to gain insight into their stability.
Outcome: The proposed results provide insights into word embedding stability in English and other languages.
A Deeper Look into Dependency-Based Word Embeddings (N18-4)

Copied to clipboard

Challenge: Word embeddings trained with dependency contexts excel at different tasks, and enhanced dependencies often improve performance.
Approach: They propose to use dependency-based word embeddings to capture semantic similarity rather than relatedness.
Outcome: The results show that word embeddings trained with Universal and Stanford dependencies excel at different tasks and that enhanced dependencies often improve performance.
Topic Sensitive Attention on Generic Corpora Corrects Sense Bias in Pretrained Embeddings (P19-1)

Copied to clipboard

Challenge: Existing methods to adapt pretrained embeddings to a large corpus are limited and do not provide sufficient quality.
Approach: They propose to use a small corpus D_T to pretrain embeddings that accurately capture the sense of words in a limited set of focused topics.
Outcome: The proposed embeddings capture the sense of words in a topic in spite of the limited size of the corpus D_T.
Quantifying Compositionality of Classic and State-of-the-Art Embeddings (2025.findings-emnlp)

Copied to clipboard

Challenge: Static word embeddings make strong claims about compositionality, but the SOTA generative models go too far in the other direction.
Approach: a new study evaluates the compositionality of word embeddings by canonical correlation analysis . strong compositional signals are observed in later training stages across data modalities .
Outcome: a new evaluation of compositional models shows that they exploit access meanings when justified . strong compositional signals are observed in later training stages and in deeper layers of the transformer-based model before a decline at the top layer.
What’s in Your Embedding, And How It Predicts Task Performance (C18-1)

Copied to clipboard

Challenge: Attempts to find a single technique for general-purpose intrinsic evaluation of word embeddings have so far not been successful.
Approach: They propose a method that quantifies interpretable characteristics of word vector neighborhoods and shows how they correlate with performance on 14 extrinsic and intrinsic task datasets.
Outcome: The proposed approach enables multi-faceted evaluation, parameter search, and generally – a more principled, hypothesis-driven approach to development of distributional semantic representations.
Sense Embeddings are also Biased – Evaluating Social Biases in Static and Contextualised Sense Embeddings (2022.acl-long)

Copied to clipboard

Challenge: Existing studies have evaluated social biases in word embeddings, but they are understudied.
Approach: They propose to evaluate the social biases in sense embeddings using a benchmark dataset for word embedders.
Outcome: The proposed measures show that even when no biases are found at word-level, there are still worrying levels of social biase at sense-level which are often ignored by the word- level bias evaluation measures.
Contextual Embeddings: When Are They Worth It? (2020.acl-main)

Copied to clipboard

Challenge: In recent years, rich contextual embeddings have enabled rapid progress on benchmarks like GLUE, but require significant computational resources during pretraining and during downstream task training and inference.
Approach: They empirically compare contextual embeddings with classic pretrained embedders and a random word embeddable with a simple baseline.
Outcome: The proposed models perform within 5 to 10% accuracy on industry-scale data.
BioReddit: Word Embeddings for User-Generated Biomedical NLP (D19-62)

Copied to clipboard

Challenge: a corpus of medical-themed posts was scrapped from Reddit to train word embeddings on downstream tasks.
Approach: They propose to train word embeddings from a corpus of medical forums from reddit scrapping posts from medical-themed subreddits.
Outcome: The proposed system outperforms embeddings trained on general purpose data or on scientific papers when applied on user-generated content.
On the Distribution of Deep Clausal Embeddings: A Large Cross-linguistic Study (P19-1)

Copied to clipboard

Challenge: Empirical evidence on the prevalence and limits of embeddings has been based on either laboratory setups or corpus data of relatively limited size.
Approach: They use large, dependency-parsed corpora to capture clausal embedding through dependency graphs and assess their distribution.
Outcome: The results show that there is no evidence for hard constraints on embedding depth . they also show that sentences with many embeddable clauses do not display a bias towards less deep embedded sentences.
Embeddings in Natural Language Processing (2020.coling-tutorials)

Copied to clipboard

Challenge: Embeddings have been a key topic of interest in NLP for the past decade . a quick warm-up introduction to NLP and why it is important to have a semantic comprehension of texts .
Approach: This tutorial will provide a high-level synthesis of the main embedding techniques in NLP . it will start with word embedds and then move to other types of embeddable vectors .
Outcome: This tutorial will provide a high-level synthesis of the main embedding techniques in NLP . it will start with word embedds and move to other types of embeddable representations .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations