Norm of Mean Contextualized Embeddings Determines their Variance (2025.coling-main)
Copied to clipboard
| Challenge: | Contextualized embeddings vary by context, even for the same token . a recent study shows a trade-off between the norm and the variance of the embedded word . |
| Approach: | They show that contextualized embeddings vary by context, even for the same token . they focus on the norm of the mean embeddment and the variance of the embeddables . |
| Outcome: | The proposed method is efficient and efficient for embeddings in sentences. |
Similar Papers
How to Dissect a Muppet: The Structure of Transformer Embedding Spaces (2022.tacl-1)
Copied to clipboard
| Challenge: | Pretrained embeddings based on the Transformer architecture have taken the NLP community by storm . a novel decomposition of Transformer output embeddables is demonstrated . |
| Approach: | They propose to decompose Transformer output embeddings into a sum of vector factors . they show multi-head attentions and feed-forwards are not equally useful in downstream applications . |
| Outcome: | The proposed method outperforms recurrent architectures on a wide variety of tasks. |
Building Static Embeddings from Contextual Ones: Is It Useful for Building Distributional Thesauri? (2022.lrec-1)
Copied to clipboard
| Challenge: | contextual language models are dominant in the field of Natural Language Processing, but they are not suitable for all uses. |
| Approach: | They propose a method for building word or type-level embeddings from contextual models . they evaluate a large set of English nouns from the perspective of extracting semantic similarity relations . |
| Outcome: | The proposed method can be used to build word or type embeddings from contextual models . it can be exploited for a wide set of English nouns, showing it can improve distributional thesauri . |
Debiasing Pre-trained Contextualised Embeddings (2021.eacl-main)
Copied to clipboard
| Challenge: | a study of contextualised word embeddings shows discriminative biases are encoded in contextualised embeddables. |
| Approach: | They propose a fine-tuning method that can be applied at token- or sentence-levels to debias pre-trained contextualised embeddings. |
| Outcome: | The proposed method can be applied at token- or sentence-levels to debias pre-trained models without requiring retrains. |
Too Much in Common: Shifting of Embeddings in Transformer Language Models and its Implications (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing studies have shown that word embeddings do not occupy a narrow cone, but rather drift in common directions. |
| Approach: | They show that anisotropy can be restored using a simple transformation of word embeddings. |
| Outcome: | The proposed model can restore anisotropy using a simple transformation. |
Frustratingly Easy Meta-Embedding – Computing Meta-Embeddings by Averaging Source Word Embeddings (N18-2)
Copied to clipboard
| Challenge: | Existing methods for producing word embeddings have shown to produce accurate meta-embeddings from pre-trained source embeddables. |
| Approach: | They propose to use arithmetic mean of two distinct word embedding sets to produce an accurate meta-embedding. |
| Outcome: | The proposed method produces meta-embeddings comparable or better than more complex methods. |
Contextual Embeddings: When Are They Worth It? (2020.acl-main)
Copied to clipboard
| Challenge: | In recent years, rich contextual embeddings have enabled rapid progress on benchmarks like GLUE, but require significant computational resources during pretraining and during downstream task training and inference. |
| Approach: | They empirically compare contextual embeddings with classic pretrained embedders and a random word embeddable with a simple baseline. |
| Outcome: | The proposed models perform within 5 to 10% accuracy on industry-scale data. |
Spying on Your Neighbors: Fine-grained Probing of Contextual Embeddings for Information about Surrounding Words (2020.acl-main)
Copied to clipboard
| Challenge: | a suite of probing tasks test contextual embeddings for encoding of information about surrounding words . authors: little is known about what information embeddables encode about the context words encode . a recent study shows that contextual embeds can be powerful for many tasks . |
| Approach: | They propose probing tasks that enable fine-grained testing of contextual embeddings . they examine popular contextual encoders and find that each encodes contextual information across tokens a little different . |
| Outcome: | The proposed probing tasks show that word embeddings encode information about words . the tests show that the encoded information is encoded across tokens with near-perfect recoverability . |
Statistical Depth for Ranking and Characterizing Transformer-Based Text Embeddings (2023.emnlp-main)
Copied to clipboard
| Challenge: | Generalized transformer-based text embedding models have produced state of the art performance results on a variety of tasks such as natural language inference (NLI) |
| Approach: | They propose a statistical depth to measure distributions of transformer-based text embeddings and an associated rank sum test to characterize distributions in synthetic and human-generated corpora. |
| Outcome: | The proposed method improves performance over baseline methods on six text classification tasks. |
Latent Positional Information is in the Self-Attention Variance of Transformer Language Models Without Positional Embeddings (2023.acl-short)
Copied to clipboard
| Challenge: | Recent research has called into question the necessity of positional embeddings in transformer language models. |
| Approach: | They propose to discard positional embeddings in transformer language models to facilitate more efficient pretraining. |
| Outcome: | The proposed model encodes strong positional information through shrinkage of self-attention variance. |
More Embeddings, Better Sequence Labelers? (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing work suggests contextual embeddings improve sequence labeling accuracy . but, there is no definite conclusion on whether concatenating different kinds of embeddables is effective . |
| Approach: | They propose a family of contextual embeddings that improves sequence labeling accuracy . they conduct extensive experiments on 3 tasks over 18 datasets and 8 languages . |
| Outcome: | The proposed family of contextual embeddings improves the accuracy of sequence labelers over non-contextual embedders. |