Frustratingly Easy Meta-Embedding – Computing Meta-Embeddings by Averaging Source Word Embeddings (N18-2)
Copied to clipboard
| Challenge: | Existing methods for producing word embeddings have shown to produce accurate meta-embeddings from pre-trained source embeddables. |
| Approach: | They propose to use arithmetic mean of two distinct word embedding sets to produce an accurate meta-embedding. |
| Outcome: | The proposed method produces meta-embeddings comparable or better than more complex methods. |
Similar Papers
Benchmarking Meta-embeddings: What Works and What Does Not (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to build meta-embeddings have been evaluated using a variety of methods and datasets, which makes it difficult to draw meaningful conclusions regarding the merits of each approach. |
| Approach: | They propose a unified framework for a fair and objective meta-embedding evaluation using intrinsic and extrinsic tasks. |
| Outcome: | The proposed framework outperforms existing methods on intrinsic and extrinsic evaluation benchmarks and outperformed existing methods. |
Dynamic Meta-Embeddings for Improved Sentence Representations (D18-1)
Copied to clipboard
| Challenge: | A sprawling literature has emerged about what word embeddings are most useful for which tasks . word embed-ding is a technique that can be used to learn word-level meaning representations for a variety of tasks. |
| Approach: | They propose a method for supervised learning of embedding ensembles that leads to state-of-the-art performance on a variety of tasks. |
| Outcome: | The proposed method leads to state-of-the-art performance on a variety of tasks. |
Learning Word Meta-Embeddings by Autoencoding (C18-1)
Copied to clipboard
| Challenge: | Existing word embeddings have shown superior performance in numerous Natural Language Processing (NLP) tasks, however, their performances vary significantly across different tasks. |
| Approach: | They propose to combine distributed word embeddings to produce more accurate and complete meta-embeddings of words. |
| Outcome: | The proposed meta-embeddings outperform the state-of-the-art in multiple tasks. |
More Embeddings, Better Sequence Labelers? (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing work suggests contextual embeddings improve sequence labeling accuracy . but, there is no definite conclusion on whether concatenating different kinds of embeddables is effective . |
| Approach: | They propose a family of contextual embeddings that improves sequence labeling accuracy . they conduct extensive experiments on 3 tasks over 18 datasets and 8 languages . |
| Outcome: | The proposed family of contextual embeddings improves the accuracy of sequence labelers over non-contextual embedders. |
Learning Efficient Task-Specific Meta-Embeddings with Word Prisms (2020.coling-main)
Copied to clipboard
| Challenge: | Word embeddings possess different lexical properties depending on the notion of context defined at training time. |
| Approach: | They introduce a meta-embedding method that learns to combine source embeddings according to the task at hand. |
| Outcome: | The proposed method improves performance on six extrinsic evaluations over other methods. |
Block-wise Word Embedding Compression Revisited: Better Weighting and Structuring (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for word embedding compression are limited . word embeds have a considerable size and need to be compressed to deploy on edge devices . |
| Approach: | They propose a block-wise low-rank approximation method for word embedding called GroupReduce . they propose 'frequency-inverse document frequency method' and a differentiable method for weighting . |
| Outcome: | The proposed algorithm more effectively finds word weights than competitors in most cases. |
Contextual Embeddings: When Are They Worth It? (2020.acl-main)
Copied to clipboard
| Challenge: | In recent years, rich contextual embeddings have enabled rapid progress on benchmarks like GLUE, but require significant computational resources during pretraining and during downstream task training and inference. |
| Approach: | They empirically compare contextual embeddings with classic pretrained embedders and a random word embeddable with a simple baseline. |
| Outcome: | The proposed models perform within 5 to 10% accuracy on industry-scale data. |
The Unreasonable Effectiveness of Random Target Embeddings for Continuous-Output Neural Machine Translation (2024.naacl-short)
Copied to clipboard
| Challenge: | Continuous-output neural machine translation models are trained to predict the continuous representation based on distances between vectors. |
| Approach: | They propose a continuous-output neural machine translation (CoNMT) approach that uses random output embeddings to outperform laboriously pre-trained models. |
| Outcome: | The proposed strategy outperforms pre-trained embeddings on large datasets and is strongest for rare words due to the geometry of their embedders. |
Text Embeddings Reveal (Almost) As Much As Text (2023.emnlp-main)
Copied to clipboard
| Challenge: | a vector database of dense text embeddings stores only the text data, not the original text . a multi-step method that iteratively corrects and re-embeds text can recover 92% of 32-token text inputs exactly. |
| Approach: | They propose a method that iteratively corrects and re-embeds text to recover 92% of 32-token text inputs exactly. |
| Outcome: | The proposed method recovers 92% of 32-token text inputs exactly. |
Addressing Noise in Multidialectal Word Embeddings (P18-2)
Copied to clipboard
| Challenge: | Dialectal Arabic (DA) is problematically noisy and lacks a large corpus of non-noisy words. |
| Approach: | They propose to use word embedding tools to maximize the informative content leveraged in each training sentence and analyze methods for representing disparate dialects in one embeddable space. |
| Outcome: | The proposed methods improve performance on low and high frequency words while preserving accuracy on low frequency forms. |