Challenge: Existing work uses sentences within the same batch as negatives, which suffers from easy negatives.
Approach: They propose to align sentence representations from different languages into a unified embedding space . they adapt MoCo to further improve the quality of alignment .
Outcome: The proposed model achieves state-of-the-art on several tasks.

Similar Papers

Improving Multi-lingual Alignment Through Soft Contrastive Learning (2024.naacl-srw)

Copied to clipboard

Challenge: Existing methods to train multi-lingual sentence embeddings ruins the mono-lingual space.
Approach: They propose a method to align multi-lingual embeddings based on similarity of sentences measured by a pre-trained mono-lingual teacher model.
Outcome: The proposed method outperforms existing multi-lingual embeddings including LaBSE on five languages and on a translation pair for Tatoeba dataset.
Dual-Alignment Pre-training for Cross-lingual Sentence Embedding (2023.acl-long)

Copied to clipboard

Challenge: Recent studies have shown that dual encoder models trained with the sentence-level translation ranking task are effective methods for cross-lingual sentence embedding.
Approach: They propose a dual-alignment pre-training framework that incorporates both sentence-level and token-level alignment.
Outcome: The proposed framework improves cross-lingual sentence embedding on three cross-linguistic benchmarks.
Semantic Alignment with Calibrated Similarity for Multilingual Sentence Embedding (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for learning semantic similarity between two English sentences have focused on one sub-task and therefore showed biased performance.
Approach: They propose a method to learn semantic similarity between two English sentences using siamese networks.
Outcome: The proposed method improves on both sub-tasks and predicts similarity scores in 14 languages.
Cross-Align: Modeling Deep Cross-lingual Interactions for Word Alignment (2022.emnlp-main)

Copied to clipboard

Challenge: Existing word alignment models capture few interactions between input sentence pairs, which severely degrades the word alignment quality.
Approach: They propose to model deep interactions between input and target sentences using a two-stage training framework to train the model.
Outcome: The proposed model achieves the state-of-the-art (SOTA) performance on four out of five language pairs.
Improving In-context Learning of Multilingual Generative Language Models with Cross-lingual Alignment (2024.naacl-long)

Copied to clipboard

Challenge: Existing studies show that multilingual generative models exhibit a strong language bias toward high-resource languages.
Approach: They propose a cross-lingual alignment framework exploiting pairs of translation sentences to improve cross-linguistic abilities.
Outcome: The proposed framework improves cross-lingual abilities and mitigates performance gap.
Cross-lingual Sentence Embedding using Multi-Task Learning (2021.emnlp-main)

Copied to clipboard

Challenge: Existing multilingual sentence embedding models require large parallel corpora to learn efficiently, limiting their scope.
Approach: They propose a sentence embedding framework based on an unsupervised loss function . they capture semantic similarity and relatedness between sentences using a multi-task loss function.
Outcome: The proposed framework outperforms state-of-the-art methods on STS, BUCC and Tatoeba benchmarks and on a monolingual benchmark.
Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment (2024.findings-naacl)

Copied to clipboard

Challenge: Current approaches to obtain cross-lingual sentence embeddings rely on pre-trained language models that implicitly align the contextual representations of similar units of sentences in different languages.
Approach: They propose a framework that explicitly aligns words between English and eight low-resource languages by using off-the-shelf word alignment models.
Outcome: The proposed framework improves on the bitext retrieval task and in high-resource languages.
Bilingual alignment transfers to multilingual alignment for unsupervised parallel text mining (2022.acl-long)

Copied to clipboard

Challenge: a model trained to align only two languages can encode multilingually more aligned representations . a dual-pivot transfer theory is proposed for bilingual training .
Approach: They propose methods for learning cross-lingual sentence representations using paired or unpaired bilingual texts.
Outcome: The proposed models reach the state of the art in unsupervised bitext mining and perform better than multilingually supervised models.
Are the Best Multilingual Document Embeddings simply Based on Sentence Embeddings? (2023.findings-eacl)

Copied to clipboard

Challenge: obtaining document embeddings at document level is challenging due to computational requirements and lack of appropriate data.
Approach: They compare methods to produce document-level representations from sentences based on LASER, LaBSE, and Sentence BERT pre-trained multilingual models.
Outcome: The proposed methods produce document-level representations from sentences in 8 languages . the results show that a clever combination of sentence embeddings is usually better than encoding the full document as a single unit.
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models.
Approach: They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key .
Outcome: The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations