Challenge: Existing studies on cosine similarity focus on the angle or correlation coefficient, but this study proposes a novel interpretation of the term word similarity.
Approach: They propose a method for selecting statistically significant axes by deriving the probability distributions that govern each component and the product of components.
Outcome: The proposed interpretation of cosine similarity is demonstrated through intuitive numerical examples and thorough numerical experiments.

Similar Papers

Exploring Intra and Inter-language Consistency in Embeddings with ICA (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that ICA can reveal universal semantic axes across languages but lack verification of consistency of independent components within and across languages.
Approach: They propose to use independent component analysis to identify independent components that are more interpretable than PCA to find universal semantic axes.
Outcome: The proposed framework ensures the reliability and universality of semantic axes.
Understanding Higher-Order Correlations Among Semantic Components in Embeddings (2024.emnlp-main)

Copied to clipboard

Challenge: Independent Component Analysis (ICA) is an effective method for visualizing and interpreting the geometric structure of embeddings.
Approach: They quantified embeddings' non-independencies using higher-order correlations and a maximum spanning tree of semantic components.
Outcome: The results provide deeper insights into embeddings through ICA.
Axis Tour: Word Tour Determines the Order of Axes in ICA-transformed Embeddings (2024.findings-emnlp)

Copied to clipboard

Challenge: Embedding is an important component in natural language processing, but interpreting high-dimensional embeddings remains challenging.
Approach: They propose a method which optimizes the order of axes in word embedding space by maximizing semantic continuity.
Outcome: The proposed method improves the clarity of the word embedding space by maximizing the semantic continuity of the axes.
Discovering Universal Geometry in Embeddings with ICA (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies have focused on achieving sparse embeddings or acquiring semantic axes, but this study focuses on the intrinsic independence present within embeddables.
Approach: They propose to use independent component analysis to extract independent semantic components from pre-trained embeddings by leveraging anisotropic information that remains after the whitening process in Principal Component Analysis.
Outcome: The proposed method reveals that embeddings can be expressed as a composition of a few interpretable axes and that these axe axe are consistent across languages, algorithms, and modalities.
Correlation Coefficients and Semantic Textual Similarity (N19-1)

Copied to clipboard

Challenge: Existing research into semantic textual similarity has focused on word embeddings . little attention has been devoted to similarity measures between word embeds - a new study shows .
Approach: They show that cosine similarity is essentially equivalent to the Pearson correlation coefficient for all common word vectors.
Outcome: The proposed model outperforms the existing model on word-level and sentence-level similarity benchmarks.
Text Similarity Estimation Based on Word Embeddings and Matrix Norms for Targeted Marketing (N19-1)

Copied to clipboard

Challenge: Existing methods to estimate document similarity based on word embeddings are mediocre . a recent study compared word and sentence embedded documents to a similarity estimate using matrix norms.
Approach: They propose to combine word embeddings with matrix norms to obtain a similarity estimate.
Outcome: The proposed method produces superior results for most of the investigated matrix norms compared to the classical cosine measure and several other similarity estimates.
Interpretable Text Embeddings and Text Similarity Explanation: A Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Text embeddings are a fundamental component in many NLP tasks, but their interpretation and explanation remain challenging.
Approach: They propose a framework for interpretable text embeddings and text similarity explanation . they characterize the main ideas, approaches, and trade-offs and discuss lessons learned .
Outcome: The proposed methods are compared with existing models and compare them with existing ones.
Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words (2022.acl-short)

Copied to clipboard

Challenge: We find that word similarities estimated by cosine over contextual embeddings are understated and trace this effect to training data frequency.
Approach: They propose to use cosine similarity to estimate word similarities in contextual embeddings to trace this effect to training data frequency.
Outcome: The proposed model underestimates similarity between frequent and low frequency words even after controlling for polysemy and other factors.
A Rank-Based Similarity Metric for Word Embeddings (P18-2)

Copied to clipboard

Challenge: Word Embeddings have become a standard for word representations, with vector cosine being the only similarity metric.
Approach: They propose to use rank-based similarity estimation metrics to measure word similarity . they find WE outperforms vector cosine in the recent outlier detection task .
Outcome: The proposed rank-based measure outperforms vector cosine in the recent outlier detection task.
Exploring Interpretability of Independent Components of Word Embeddings with Automated Word Intruder Test (2024.lrec-main)

Copied to clipboard

Challenge: Independent Component Analysis (ICA) is an algorithm for finding separate sources in a mixed signal.
Approach: They propose to use ICA to analyze word embeddings to quantify interpretability . they propose to automate word intruder test to quantify the components .
Outcome: The proposed algorithm can be used to find semantic features of words . it can be combined to find words that have features associated with the components .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations