The Word Analogy Testing Caveat (N18-2)

Copied to clipboard

Challenge: a number of word analogy tests are used to evaluate word embeddings . word embeds are used as a proxy for semantics and syntax à la Harris .
Approach: They propose to use word embeddings as a proxy for distributional similarity . they propose to apply a transfer learning approach to word embeds to improve performance .
Outcome: The proposed method improves performance across a wide range of NLP tasks.

Similar Papers

Paraphrases do not explain word analogies (2021.eacl-main)

Copied to clipboard

Challenge: Several attempts have been made to explain distributional word embeddings as linguistic regularities as directions.
Approach: They propose to use an analogy to explain why linguistic regularities should hold in distributional word embeddings.
Outcome: The proposed explanation does not hold empirically.
On the Correlation of Word Embedding Evaluation Metrics (2020.lrec-1)

Copied to clipboard

Challenge: Word embeddings are geometrical representations of word paradigmatics and syntagmatics.
Approach: They propose to investigate evaluation metrics on various datasets to find correlations . they propose a fast solution to select the best word embeddings among many others .
Outcome: The proposed method could be used to select the best word embeddings among many others.
A Deeper Look into Dependency-Based Word Embeddings (N18-4)

Copied to clipboard

Challenge: Word embeddings trained with dependency contexts excel at different tasks, and enhanced dependencies often improve performance.
Approach: They propose to use dependency-based word embeddings to capture semantic similarity rather than relatedness.
Outcome: The results show that word embeddings trained with Universal and Stanford dependencies excel at different tasks and that enhanced dependencies often improve performance.
Robustness and Reliability of Gender Bias Assessment in Word Embeddings: The Role of Base Pairs (2020.aacl-main)

Copied to clipboard

Challenge: Existing methods to quantify gender bias in word embeddings are not robust and cannot identify common types of bias.
Approach: They propose to quantify gender bias by using cosine similarity to a pair of gender words and using analogies.
Outcome: The proposed methods are not robust and cannot identify common types of bias, while analogies are unsuitable indicators.
What’s in Your Embedding, And How It Predicts Task Performance (C18-1)

Copied to clipboard

Challenge: Attempts to find a single technique for general-purpose intrinsic evaluation of word embeddings have so far not been successful.
Approach: They propose a method that quantifies interpretable characteristics of word vector neighborhoods and shows how they correlate with performance on 14 extrinsic and intrinsic task datasets.
Outcome: The proposed approach enables multi-faceted evaluation, parameter search, and generally – a more principled, hypothesis-driven approach to development of distributional semantic representations.
Analyzing the Surprising Variability in Word Embedding Stability Across Languages (2021.emnlp-main)

Copied to clipboard

Challenge: Word embeddings are powerful representations that form the foundation of many natural language processing architectures.
Approach: They explore word embedding stability in a wide range of languages to gain insight into their stability.
Outcome: The proposed results provide insights into word embedding stability in English and other languages.
Conditional Word Embedding and Hypothesis Testing via Bayes-by-Backprop (D18-1)

Copied to clipboard

Challenge: Whether word's meaning varies across contexts has become a major focus of research in recent years.
Approach: They propose a word embedding model that incorporates document covariates to estimate conditional word embeds.
Outcome: The proposed model estimates word embedding distributions based on document covariates . if word embeds are statistically significant, hypothesis tests can be performed .
Factors Influencing the Surprising Instability of Word Embeddings (N18-1)

Copied to clipboard

Challenge: Word embeddings are low-dimensional, dense vector representations that capture semantic properties of words.
Approach: They examine the stability of word embeddings by examining their properties and analyzing their effects on downstream tasks.
Outcome: The results show that even high frequency words exhibit substantial instability, which can have implications for downstream tasks.
Sentence Analogies: Linguistic Regularities in Sentence Embeddings (2020.coling-main)

Copied to clipboard

Challenge: Word vectors are often evaluated by assessing to what degree they exhibit regularities with regard to relationships considered in word analogies.
Approach: They propose a number of schemes to induce evaluation data based on lexical analogy data as well as semantic relationships between sentences.
Outcome: The proposed models reflect regularities in lexical analogies and semantic relationships between sentences.
When do Word Embeddings Accurately Reflect Surveys on our Beliefs About People? (2020.acl-main)

Copied to clipboard

Challenge: a study of word embeddings shows that social biases are more accurate than survey data for some dimensions of meaning.
Approach: a new study investigates the extent to which word embeddings accurately reflect biases . they find that biased word embeds mirror survey data across 17 dimensions of social meaning .
Outcome: a new study shows that word embeddings accurately reflect biases on average across dimensions of social meaning . biased embedders are more reflective of survey data for some dimensions of meaning than others, the study finds .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations