Discovering Universal Geometry in Embeddings with ICA (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies have focused on achieving sparse embeddings or acquiring semantic axes, but this study focuses on the intrinsic independence present within embeddables.
Approach: They propose to use independent component analysis to extract independent semantic components from pre-trained embeddings by leveraging anisotropic information that remains after the whitening process in Principal Component Analysis.
Outcome: The proposed method reveals that embeddings can be expressed as a composition of a few interpretable axes and that these axe axe are consistent across languages, algorithms, and modalities.

Similar Papers

Exploring Intra and Inter-language Consistency in Embeddings with ICA (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that ICA can reveal universal semantic axes across languages but lack verification of consistency of independent components within and across languages.
Approach: They propose to use independent component analysis to identify independent components that are more interpretable than PCA to find universal semantic axes.
Outcome: The proposed framework ensures the reliability and universality of semantic axes.
Axis Tour: Word Tour Determines the Order of Axes in ICA-transformed Embeddings (2024.findings-emnlp)

Copied to clipboard

Challenge: Embedding is an important component in natural language processing, but interpreting high-dimensional embeddings remains challenging.
Approach: They propose a method which optimizes the order of axes in word embedding space by maximizing semantic continuity.
Outcome: The proposed method improves the clarity of the word embedding space by maximizing the semantic continuity of the axes.
Understanding Higher-Order Correlations Among Semantic Components in Embeddings (2024.emnlp-main)

Copied to clipboard

Challenge: Independent Component Analysis (ICA) is an effective method for visualizing and interpreting the geometric structure of embeddings.
Approach: They quantified embeddings' non-independencies using higher-order correlations and a maximum spanning tree of semantic components.
Outcome: The results provide deeper insights into embeddings through ICA.
Autoencoding Improves Pre-trained Word Embeddings (2020.coling-main)

Copied to clipboard

Challenge: Existing work has shown that word embeddings are distributed in a narrow cone and that centering and projection can improve the accuracy of pre-trained word embeds without requiring additional training data.
Approach: They propose to remove the top principal components from pre-trained word embeddings and center and project them onto principal component vectors to reinstate isotropy in the embeddable space.
Outcome: The proposed method is equivalent to applying a linear autoencoder to minimize the squared L2 reconstruction error.
Revisiting Cosine Similarity via Normalized ICA-transformed Embeddings (2025.coling-main)

Copied to clipboard

Challenge: Existing studies on cosine similarity focus on the angle or correlation coefficient, but this study proposes a novel interpretation of the term word similarity.
Approach: They propose a method for selecting statistically significant axes by deriving the probability distributions that govern each component and the product of components.
Outcome: The proposed interpretation of cosine similarity is demonstrated through intuitive numerical examples and thorough numerical experiments.
Interpretable Text Embeddings and Text Similarity Explanation: A Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Text embeddings are a fundamental component in many NLP tasks, but their interpretation and explanation remain challenging.
Approach: They propose a framework for interpretable text embeddings and text similarity explanation . they characterize the main ideas, approaches, and trade-offs and discuss lessons learned .
Outcome: The proposed methods are compared with existing models and compare them with existing ones.
Semantic Geometry of Sentence Embeddings (2025.findings-emnlp)

Copied to clipboard

Challenge: Sentence embeddings are central to natural language processing, but their internal features are not interpretable and users lack fine-grained control for downstream tasks.
Approach: They propose a formal framework to characterize the organization of features in sentence embeddings . they show how they can be composed to capture richer semantic structures .
Outcome: The proposed method can be used to capture richer semantic structures.
Describing Sets of Images with Textual-PCA (2022.findings-emnlp)

Copied to clipboard

Challenge: a new method to describe images using a common theme is needed to describe the images . a grammatical phrase is not sufficient to describe an image set, since captioning engines are not general enough.
Approach: They propose a method to capture attributes of images and variations within a set . they use a pretrained vision-language model to generate a centroid phrase with the largest average similarity .
Outcome: The proposed method captures the essence of image sets and describes them in a semantically meaningful way . it is easy for humans to identify and describe a common theme, but it is not generic enough .
What’s in Your Embedding, And How It Predicts Task Performance (C18-1)

Copied to clipboard

Challenge: Attempts to find a single technique for general-purpose intrinsic evaluation of word embeddings have so far not been successful.
Approach: They propose a method that quantifies interpretable characteristics of word vector neighborhoods and shows how they correlate with performance on 14 extrinsic and intrinsic task datasets.
Outcome: The proposed approach enables multi-faceted evaluation, parameter search, and generally – a more principled, hypothesis-driven approach to development of distributional semantic representations.
Parallax: Visualizing and Understanding the Semantics of Embedding Spaces via Algebraic Formulae (P19-3)

Copied to clipboard

Challenge: Embeddings are a fundamental component of many modern machine learning and natural language processing models.
Approach: They propose a tool for visualizing embedding spaces using parametric projections . they demonstrate the power of Parallax and propose % task-oriented approach .
Outcome: The proposed tool is based on two-dimensional projections without interpretable semantics . it enhances interpretability and allows for more fine-grained analysis .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations