| Challenge: | Monolingual dictionaries are widespread and semantically rich resources. |
| Approach: | They propose a model that learns to compute word embeddings by processing dictionary definitions and trying to reconstruct them. |
| Outcome: | The proposed model shows strong performance when trained exclusively on dictionary data and generalizes in one shot. |
Similar Papers
Learning Word Meta-Embeddings by Autoencoding (C18-1)
Copied to clipboard
| Challenge: | Existing word embeddings have shown superior performance in numerous Natural Language Processing (NLP) tasks, however, their performances vary significantly across different tasks. |
| Approach: | They propose to combine distributed word embeddings to produce more accurate and complete meta-embeddings of words. |
| Outcome: | The proposed meta-embeddings outperform the state-of-the-art in multiple tasks. |
Lacking the Embedding of a Word? Look it up into a Traditional Dictionary (2022.findings-acl)
Copied to clipboard
Elena Sofia Ruzzetti, Leonardo Ranaldi, Michele Mastromattei, Francesca Fallucchi, Noemi Scarpato, Fabio Massimo Zanzotto
| Challenge: | Word embeddings are powerful dictionaries, but they fail to give sense to rare words . a large body of research is devoted to devising ways to capture word meaning . |
| Approach: | They propose to use definitions retrieved from traditional dictionaries to build word embeddings for rare words. |
| Outcome: | The proposed methods outperform state-of-the-art methods for embeddings of unknown words . the proposed methods significantly outperformed the BERT method for OOV words compared to the proposed method . |
Dictionary-based Debiasing of Pre-trained Word Embeddings (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing methods for learning word embeddings using dictionaries do not require access to training resources or knowledge regarding the word embeds used. |
| Approach: | They propose a method for debiasing pre-trained word embeddings using dictionaries . they learn constraints that must be satisfied by unbiased word embeds from dictionary definitions . |
| Outcome: | The proposed method removes unfair biases encoded in pre-trained word embeddings while preserving useful semantics. |
Automatically Generated Definitions and their utility for Modeling Word Meaning (2024.emnlp-main)
Copied to clipboard
| Challenge: | Modern language models generate semantic representations for words based on context and context based models. |
| Approach: | They propose to use dictionary-like sense definitions to generate sentence embeddings . they evaluate the quality of the generated definitions on existing English benchmarks based on the results of their study . |
| Outcome: | The proposed model sets new state-of-the-art results on lexical semantics tasks compared to baselines . |
HG2Vec: Improved Word Embeddings from Dictionary and Thesaurus Based Heterogeneous Graph (2022.coling-1)
Copied to clipboard
| Challenge: | Existing models that learn word embeddings rely on a large corpus of data . however, these models require massive time and space for data pre-processing and training . |
| Approach: | They propose a model that learns word embeddings utilizing only dictionaries and thesauri . they exploit a new context-focused loss model that models transitive relationships between word pairs . |
| Outcome: | The proposed model reaches the state-of-art on multiple word similarity and relatedness benchmarks. |
Learning Bias-reduced Word Embeddings Using Dictionary Definitions (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing word embeddings have undesirable gender, racial, and religious biases . DD-GloVe is a train-time debiasing algorithm that uses dictionary definitions based on word definitions. |
| Approach: | They propose a dictionary-guided loss function that encourages word embeddings to be similar to their relatively neutral dictionary definition representations. |
| Outcome: | The proposed algorithm can learn word embeddings by leveraging dictionary definitions. |
Leveraging a Bilingual Dictionary to Learn Wolastoqey Word Representations (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing word embeddings for lowresource languages require large corpora of running text to learn high quality representations. |
| Approach: | They leverage a bilingual dictionary to learn Wolastoqey word embeddings by encoding their corresponding English definitions into vector representations using pretrained English word and sequence representation models. |
| Outcome: | The proposed model outperforms baseline models without language-specific training or fine-tuning. |
A Simple Approach to Learning Unsupervised Multilingual Embeddings (2020.emnlp-main)
Copied to clipboard
| Challenge: | Recent work on unsupervised cross-lingual embeddings in the bilingual setting has given the impetus to learning a shared embeddable space for several languages. |
| Approach: | They propose to solve two sub-problems together to learn a shared embedding space for several languages. |
| Outcome: | The proposed approach outperforms existing methods in bilingual lexicon induction, cross-lingual word similarity, multilingual document classification, and multilingual dependency parsing tasks. |
A Unified Model for Reverse Dictionary and Definition Modelling (2022.aacl-short)
Copied to clipboard
| Challenge: | Using neural networks, we argue that both tasks can be learned and dealt with concurrently, based on the intuition that a word and its definition share the same meaning. |
| Approach: | They build a dual-way neural dictionary to retrieve words given definitions and produce definitions for queried words. |
| Outcome: | The proposed model achieves high scores on previous benchmarks without extra resources. |
Building Static Embeddings from Contextual Ones: Is It Useful for Building Distributional Thesauri? (2022.lrec-1)
Copied to clipboard
| Challenge: | contextual language models are dominant in the field of Natural Language Processing, but they are not suitable for all uses. |
| Approach: | They propose a method for building word or type-level embeddings from contextual models . they evaluate a large set of English nouns from the perspective of extracting semantic similarity relations . |
| Outcome: | The proposed method can be used to build word or type embeddings from contextual models . it can be exploited for a wide set of English nouns, showing it can improve distributional thesauri . |