Aligning Vector-spaces with Noisy Supervised Lexicon (N19-1)

Copied to clipboard

Challenge: Current approaches to learning to translate between two vector spaces assume that the lexicon defines alignment pairs is noise-free.
Approach: They propose a model that accounts for noisy pairs and propose supervised learning problems for this problem.
Outcome: The proposed model significantly improves translation accuracy on bilingual word embedding translation and mapping between diachronic embeddable spaces.

Similar Papers

Bilingual Lexicon Induction via Unsupervised Bitext Construction and Word Alignment (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for bilingual lexicon induction are linear and require simplifying assumptions.
Approach: They propose methods that combine unsupervised bitext mining and unsupervised word alignment to produce higher quality lexicons.
Outcome: The proposed method outperforms the state-of-the-art on the BUCC 2020 task by 14 F1 points . further analysis suggests they are comparable quality .
Deep Generative Model for Joint Alignment and Word Representation (N18-1)

Copied to clipboard

Challenge: EmbedAlign model embeds words in their complete observed context and learns by marginalisation of latent lexical alignments.
Approach: They exploit translation as a distributional context and embed words as posterior probability densities, rather than point estimates, which allows them to compare words in context using a measure of overlap between distributions.
Outcome: The proposed model performs on a range of lexical semantics tasks and achieves competitive results on benchmarks including natural language inference, paraphrasing, and text similarity.
Supervised and Nonlinear Alignment of Two Embedding Spaces for Dictionary Induction in Low Resourced Languages (D19-1)

Copied to clipboard

Challenge: Existing methods for mapping monolingual word embeddings into another are based on anchor points and unsupervised methods are more adversarial.
Approach: They propose a noise-tolerant piecewise linear technique to learn a non-linear mapping between two monolingual word embedding vector spaces.
Outcome: The proposed method outperforms the state-of-the-art in lower resourced settings with an average of 3.7% improvement of precision @10 across 14 mostly low resourced languages.
A Simple Approach to Learning Unsupervised Multilingual Embeddings (2020.emnlp-main)

Copied to clipboard

Challenge: Recent work on unsupervised cross-lingual embeddings in the bilingual setting has given the impetus to learning a shared embeddable space for several languages.
Approach: They propose to solve two sub-problems together to learn a shared embedding space for several languages.
Outcome: The proposed approach outperforms existing methods in bilingual lexicon induction, cross-lingual word similarity, multilingual document classification, and multilingual dependency parsing tasks.
How Lexical is Bilingual Lexicon Induction? (2024.findings-naacl)

Copied to clipboard

Challenge: lexical variation and low-resource settings make it difficult to learn in low-level settings.
Approach: They propose to incorporate additional lexical information into the retrieve-and-rank approach to improve lexicon induction.
Outcome: The proposed approach improves on XLING by an average of 2% across all language pairs.
Geometry-aware domain adaptation for unsupervised alignment of word embeddings (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for learning bilingual word embeddings have been used in natural language processing.
Approach: They propose a manifold based geometric approach for learning unsupervised alignment of word embeddings between the source and target languages.
Outcome: The proposed approach outperforms state-of-the-art optimal transport based approach on bilingual lexicon induction task across several language pairs.
Are All Good Word Vector Spaces Isomorphic? (2020.emnlp-main)

Copied to clipboard

Challenge: Existing algorithms for aligning cross-lingual word vector spaces assume that vector spaces are approximately isomorphic.
Approach: They propose to find out whether non-isomorphism is also crucially a sign of degenerate word vector spaces.
Outcome: The proposed method performs poorly on non-isomorphic spaces, but it is not . it is also crucially a sign of degenerate word vector spaces, the authors show .
Improving Cross-Lingual Word Embeddings by Meeting in the Middle (D18-1)

Copied to clipboard

Challenge: Cross-lingual word embeddings are becoming increasingly important in multilingual NLP.
Approach: They propose to apply an additional transformation after initial alignment to align two disjoint monolingual vector spaces.
Outcome: The proposed approach outperforms state-of-the-art models in monolingual and cross-lingual evaluation tasks.
DM-BLI: Dynamic Multiple Subspaces Alignment for Unsupervised Bilingual Lexicon Induction (2024.acl-long)

Copied to clipboard

Challenge: Existing approaches to unsupervised bilingual lexicon induction (BLI) fail on distant or low-resource language pairs, achieving less than half the performance observed in rich-resourced languages.
Approach: They propose a framework for unsupervised bilingual lexicon induction that uses multiple subspace alignments instead of a single mapping.
Outcome: Experiments on 12 language pairs show that the proposed framework improves language alignment by utilizing multiple subspace alignments instead of a single mapping.
On the Robustness of Unsupervised and Semi-supervised Cross-lingual Word Embedding Learning (2020.lrec-1)

Copied to clipboard

Challenge: Cross-lingual word embeddings are vector representations of words in different languages where words with similar meaning are represented by similar vectors, regardless of the language.
Approach: They propose to evaluate multiple cross-lingual word embedding models and compare their strengths and limitations to evaluate their effectiveness.
Outcome: The proposed models perform well with noisy text and language pairs with major differences.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations