| Challenge: | Current approaches to learning to translate between two vector spaces assume that the lexicon defines alignment pairs is noise-free. |
| Approach: | They propose a model that accounts for noisy pairs and propose supervised learning problems for this problem. |
| Outcome: | The proposed model significantly improves translation accuracy on bilingual word embedding translation and mapping between diachronic embeddable spaces. |
Similar Papers
Bilingual Lexicon Induction via Unsupervised Bitext Construction and Word Alignment (2021.acl-long)
Copied to clipboard
| Challenge: | Existing methods for bilingual lexicon induction are linear and require simplifying assumptions. |
| Approach: | They propose methods that combine unsupervised bitext mining and unsupervised word alignment to produce higher quality lexicons. |
| Outcome: | The proposed method outperforms the state-of-the-art on the BUCC 2020 task by 14 F1 points . further analysis suggests they are comparable quality . |
Deep Generative Model for Joint Alignment and Word Representation (N18-1)
Copied to clipboard
| Challenge: | EmbedAlign model embeds words in their complete observed context and learns by marginalisation of latent lexical alignments. |
| Approach: | They exploit translation as a distributional context and embed words as posterior probability densities, rather than point estimates, which allows them to compare words in context using a measure of overlap between distributions. |
| Outcome: | The proposed model performs on a range of lexical semantics tasks and achieves competitive results on benchmarks including natural language inference, paraphrasing, and text similarity. |
Supervised and Nonlinear Alignment of Two Embedding Spaces for Dictionary Induction in Low Resourced Languages (D19-1)
Copied to clipboard
| Challenge: | Existing methods for mapping monolingual word embeddings into another are based on anchor points and unsupervised methods are more adversarial. |
| Approach: | They propose a noise-tolerant piecewise linear technique to learn a non-linear mapping between two monolingual word embedding vector spaces. |
| Outcome: | The proposed method outperforms the state-of-the-art in lower resourced settings with an average of 3.7% improvement of precision @10 across 14 mostly low resourced languages. |
A Simple Approach to Learning Unsupervised Multilingual Embeddings (2020.emnlp-main)
Copied to clipboard
| Challenge: | Recent work on unsupervised cross-lingual embeddings in the bilingual setting has given the impetus to learning a shared embeddable space for several languages. |
| Approach: | They propose to solve two sub-problems together to learn a shared embedding space for several languages. |
| Outcome: | The proposed approach outperforms existing methods in bilingual lexicon induction, cross-lingual word similarity, multilingual document classification, and multilingual dependency parsing tasks. |
How Lexical is Bilingual Lexicon Induction? (2024.findings-naacl)
Copied to clipboard
| Challenge: | lexical variation and low-resource settings make it difficult to learn in low-level settings. |
| Approach: | They propose to incorporate additional lexical information into the retrieve-and-rank approach to improve lexicon induction. |
| Outcome: | The proposed approach improves on XLING by an average of 2% across all language pairs. |
Geometry-aware domain adaptation for unsupervised alignment of word embeddings (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods for learning bilingual word embeddings have been used in natural language processing. |
| Approach: | They propose a manifold based geometric approach for learning unsupervised alignment of word embeddings between the source and target languages. |
| Outcome: | The proposed approach outperforms state-of-the-art optimal transport based approach on bilingual lexicon induction task across several language pairs. |
Are All Good Word Vector Spaces Isomorphic? (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing algorithms for aligning cross-lingual word vector spaces assume that vector spaces are approximately isomorphic. |
| Approach: | They propose to find out whether non-isomorphism is also crucially a sign of degenerate word vector spaces. |
| Outcome: | The proposed method performs poorly on non-isomorphic spaces, but it is not . it is also crucially a sign of degenerate word vector spaces, the authors show . |
Improving Cross-Lingual Word Embeddings by Meeting in the Middle (D18-1)
Copied to clipboard
| Challenge: | Cross-lingual word embeddings are becoming increasingly important in multilingual NLP. |
| Approach: | They propose to apply an additional transformation after initial alignment to align two disjoint monolingual vector spaces. |
| Outcome: | The proposed approach outperforms state-of-the-art models in monolingual and cross-lingual evaluation tasks. |
DM-BLI: Dynamic Multiple Subspaces Alignment for Unsupervised Bilingual Lexicon Induction (2024.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to unsupervised bilingual lexicon induction (BLI) fail on distant or low-resource language pairs, achieving less than half the performance observed in rich-resourced languages. |
| Approach: | They propose a framework for unsupervised bilingual lexicon induction that uses multiple subspace alignments instead of a single mapping. |
| Outcome: | Experiments on 12 language pairs show that the proposed framework improves language alignment by utilizing multiple subspace alignments instead of a single mapping. |
On the Robustness of Unsupervised and Semi-supervised Cross-lingual Word Embedding Learning (2020.lrec-1)
Copied to clipboard
| Challenge: | Cross-lingual word embeddings are vector representations of words in different languages where words with similar meaning are represented by similar vectors, regardless of the language. |
| Approach: | They propose to evaluate multiple cross-lingual word embedding models and compare their strengths and limitations to evaluate their effectiveness. |
| Outcome: | The proposed models perform well with noisy text and language pairs with major differences. |