Challenge: Multilingual pretrained language models (mPLMs) have shown their effectiveness in multilingual word alignment induction, but these methods usually start from mBERT or XLM-R.
Approach: They propose to fine tune multilingual sentence Transformer LaBSE for alignment induction using parallel corpus and a parallel corpora model.
Outcome: The proposed model outperforms existing models on seven language pairs and achieves new state-of-the-art on zero-shot language pairs.

Similar Papers

Word Alignment by Fine-tuning Embeddings on Parallel Corpora (2021.eacl-main)

Copied to clipboard

Challenge: Existing work on word alignment has focused on unsupervised learning on parallel text.
Approach: They propose to combine pre-trained contextualized word embeddings with multilingually trained language models to achieve competitive results on word alignment tasks.
Outcome: The proposed model outperforms state-of-the-art models on five language pairs and can train multilingual word aligners that can obtain robust performance on different language pairs.
Multilingual BERT Post-Pretraining Alignment (2021.naacl-main)

Copied to clipboard

Challenge: Recent work improves on the success of monolingual pretrained language models by adding cross-lingual tasks that always involve English.
Approach: They propose a method to align multilingual contextual embeddings as a post-pretraining step for improved cross-lingual transferability of pretrained language models.
Outcome: The proposed model outperforms XLM-R_Base on translation-train tasks while using less parallel data and fewer parameters.
A Discriminative Neural Model for Cross-Lingual Word Alignment (D19-1)

Copied to clipboard

Challenge: a novel word alignment model for machine translation has been developed for a number of languages . explicit word-to-word alignments have largely been lost in neural MT systems .
Approach: They propose a discriminative word alignment model which integrates into a Transformer-based machine translation model.
Outcome: The proposed model performs better on Chinese and Arabic alignments than standard models.
Meeting the Needs of Low-Resource Languages: The Value of Automatic Alignments via Pretrained Models (2023.eacl-main)

Copied to clipboard

Challenge: Large multilingual models have inspired a new class of word alignment methods, which work well for pretraining languages.
Approach: They propose to use transformer-based word alignment methods to extract alignments from massive pretrained models.
Outcome: The proposed methods outperform traditional methods for languages unseen to pretraining models, and are competitive with each other.
Exploring the Relationship between Alignment and Cross-lingual Transfer in Multilingual Transformers (2023.findings-acl)

Copied to clipboard

Challenge: despite lack of explicit cross-lingual training data, multilingual models can achieve cross-linguistic transfer.
Approach: They find alignment is significantly correlated with cross-lingual transfer . they advocate for further research on realignment methods for smaller models .
Outcome: The proposed method outperforms XLM-R Large in POS-tagging between English and Arabic by +15.8 accuracy.
Are the Best Multilingual Document Embeddings simply Based on Sentence Embeddings? (2023.findings-eacl)

Copied to clipboard

Challenge: obtaining document embeddings at document level is challenging due to computational requirements and lack of appropriate data.
Approach: They compare methods to produce document-level representations from sentences based on LASER, LaBSE, and Sentence BERT pre-trained multilingual models.
Outcome: The proposed methods produce document-level representations from sentences in 8 languages . the results show that a clever combination of sentence embeddings is usually better than encoding the full document as a single unit.
SimAlign: High Quality Word Alignments Without Parallel Training Data Using Static and Contextualized Embeddings (2020.findings-emnlp)

Copied to clipboard

Challenge: Word alignments are useful for statistical and neural machine translation (NMT) and cross-lingual annotation projection.
Approach: They propose to leverage multilingual word embeddings for word alignment.
Outcome: The proposed methods perform better for four languages and comparable for two languages than traditional statistical aligners even with abundant parallel data.
AlignX: Advancing Multilingual Large Language Models with Multilingual Representation Alignment (2025.emnlp-main)

Copied to clipboard

Challenge: Multilingual large language models (LLMs) possess impressive multilingual understanding and generation capabilities, but performance and cross-lingual alignment often lag for non-dominant languages.
Approach: They propose a representation-level framework to enhance multilingual performance of pre-trained LLMs by integrating multilingual semantic alignment and language feature integration.
Outcome: The proposed framework improves multilingual capability of pre-trained LLMs by bringing representations closer and improving cross-lingual alignment.
Modular Sentence Encoders: Separating Language Specialization from Cross-Lingual Alignment (2025.acl-long)

Copied to clipboard

Challenge: Multilingual sentence encoders are often trained to map sentences from different languages into a shared semantic vector space . cross-lingual alignment training distorts optimal monolingual structure of semantic spaces of individual languages . a modular solution can be used for cross-linguistic tasks such as cross-language semantic similarity and zero-shot transfer .
Approach: They propose a modular training system that embeds sentences from different languages into a shared semantic vector space.
Outcome: The proposed solution achieves better performance across all tasks compared to monolithic models.
Aligning Multilingual Embeddings for Improved Code-switched Natural Language Understanding (2022.coling-1)

Copied to clipboard

Challenge: a recent study has shown that multilingual models can be effective on monolingual data but need additional training to work well with code-switched text.
Approach: They propose to train multilingual models with alignment objectives using parallel text . they find such an explicit alignment step improves performance on code-switched NLP tasks .
Outcome: The proposed model improves on Hindi-English Sentiment Analysis, Named Entity Recognition and Question Answering tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations