Challenge: Existing word embedding algorithms make a strong assumption that words are semantically related only if they co-occur locally within a window of fixed size.
Approach: They propose a graph-based word embedding method that relies on locality to capture the semantic association between words that co-occur frequently but non-locally within documents.
Outcome: The proposed method outperforms word2vec and glove on a range of different tasks, such as predicting word-pair similarity, word analogy and concept categorization.

Similar Papers

Word Mover’s Embedding: From Word2Vec to Document Embedding (D18-1)

Copied to clipboard

Challenge: Recent work has demonstrated that Word Mover’s Distance (WMD) that aligns semantically similar words yields unprecedented KNN classification accuracy.
Approach: They propose a Word Mover’s Distance (WMD) method that aligns semantically similar words to generate unsupervised sentences or documents embeddings.
Outcome: The proposed method consistently outperforms state-of-the-art techniques on 9 benchmark text classification datasets and 22 textual similarity tasks.
Deconstructing word embedding algorithms (2020.emnlp-main)

Copied to clipboard

Challenge: Word embeddings are reliable feature representations of words used in many NLP tasks today.
Approach: They propose to deconstruct Word2vec, GloVe and others into a common form . they propose to generalize several word embedding algorithms into . a low rank embedder framework is proposed to generalise the algorithms into one common form.
Outcome: The proposed framework can be used to make word embeddings more performant.
Domain-Specific Word Embeddings with Structure Prediction (2023.tacl-1)

Copied to clipboard

Challenge: Current word embedding methods do not provide a way to use or predict information on structure between sub-corpora, time or domain.
Approach: They propose a word embedding method that provides general word representations for the whole corpus, domain-specific representations and embeddable alignment simultaneously.
Outcome: The proposed method provides better performance than baselines on a dataset of science and philosophy articles.
attr2vec: Jointly Learning Word and Contextual Attribute Embeddings with Factorization Machines (N18-1)

Copied to clipboard

Challenge: popular word embeddings are used to learn vector representations from the context of words.
Approach: They propose a framework for jointly learning embeddings for words and contextual attributes based on factorization machines.
Outcome: The proposed framework improves on a text classification task compared to learning embeddings independently.
Querying Word Embeddings for Similarity and Relatedness (N18-1)

Copied to clipboard

Challenge: Word2Vec embeddings have become popular representations of word meaning . similarity between two words is often assumed to be a direction-less measure, whereas relatedness is inherently directional.
Approach: They propose to use word embeddings to predict asymmetric association between words from a dataset of production norms to generate thematically related words.
Outcome: The proposed model predicts asymmetric association between words from a recently published dataset of production norms.
Embeddings in Natural Language Processing (2020.coling-tutorials)

Copied to clipboard

Challenge: Embeddings have been a key topic of interest in NLP for the past decade . a quick warm-up introduction to NLP and why it is important to have a semantic comprehension of texts .
Approach: This tutorial will provide a high-level synthesis of the main embedding techniques in NLP . it will start with word embedds and then move to other types of embeddable vectors .
Outcome: This tutorial will provide a high-level synthesis of the main embedding techniques in NLP . it will start with word embedds and move to other types of embeddable representations .
Dirichlet-Smoothed Word Embeddings for Low-Resource Settings (2020.lrec-1)

Copied to clipboard

Challenge: Existing count-based word embeddings are superseded by machine-learning methods like word2vec and GloVe, but in many settings there is not much text data available.
Approach: They propose to use positive pointwise mutual information (PPMI) weighted co-occurrence matrices to compute word embeddings from a corpus using large amounts of text data.
Outcome: The proposed method outperforms word2vec and the state-of-the-art for low-resource settings and obtains competitive results for Maltese and Luxembourgish.
How to represent a word and predict it, too: Improving tied architectures for language modelling (D18-1)

Copied to clipboard

Challenge: Recent state-of-the-art models use word embeddings as input and output mappings instead of tied models.
Approach: They propose to decouple hidden state from word embedding prediction . they extend their proposed modification to word2vec models .
Outcome: The proposed architectures achieve comparable or better results compared to previous models without tying . the proposed architecture reduces parameters, enabling more compact models and faster learning.
Embedding Words in Non-Vector Space with Unsupervised Graph Learning (2020.emnlp-main)

Copied to clipboard

Challenge: GraphGlove is an unsupervised graph word representations that are learned end-to-end.
Approach: They propose a method to learn weighted graph word representations end-to-end using a weighteable weighte . they adopt a hierarchical graph representation method and modify the GloVe training algorithm to learn graph representations.
Outcome: The proposed method outperforms vector-based methods on word similarity and analogy tasks.
Enhanced Word Representations for Bridging Anaphora Resolution (N18-2)

Copied to clipboard

Challenge: Existing word representations do not capture semantic similarity for bridging anaphora resolution.
Approach: They propose to use word embeddings to capture semantic similarity by exploring syntactic structure of noun phrases.
Outcome: The proposed model achieves 30% of accuracy for bridging anaphora resolution on ISNotes corpus.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations