Challenge: Document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other.
Approach: They propose an unsupervised scoring function that leverages cross-lingual sentence embeddings to compute the semantic distance between documents in different languages.
Outcome: The proposed scoring function outperforms baseline methods on high-resource language pairs, 15% on mid-resourced language pairs and 22% on low-resourcing language pairs.

Similar Papers

CCAligned: A Massive Collection of Cross-Lingual Web-Document Pairs (2020.emnlp-main)

Copied to clipboard

Challenge: Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other.
Approach: They exploit the signals embedded in URLs to label web documents at scale with an average precision of 94.5% across different language pairs.
Outcome: The proposed method can label documents at 94.5% across languages with high precision . the proposed method is useful for low-resource languages with limited resources .
Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment (2024.findings-naacl)

Copied to clipboard

Challenge: Current approaches to obtain cross-lingual sentence embeddings rely on pre-trained language models that implicitly align the contextual representations of similar units of sentences in different languages.
Approach: They propose a framework that explicitly aligns words between English and eight low-resource languages by using off-the-shelf word alignment models.
Outcome: The proposed framework improves on the bitext retrieval task and in high-resource languages.
BiMax: Bidirectional MaxSim Score for Document-Level Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: Document alignment is necessary for the hierarchical mining of documents across source and target languages.
Approach: They propose a cross-lingual Bidirectional Maxsim score for computing doc-to-doc similarity.
Outcome: The proposed method achieves accuracy comparable to OT with an approximate 100-fold speed increase.
Cross-Align: Modeling Deep Cross-lingual Interactions for Word Alignment (2022.emnlp-main)

Copied to clipboard

Challenge: Existing word alignment models capture few interactions between input sentence pairs, which severely degrades the word alignment quality.
Approach: They propose to model deep interactions between input and target sentences using a two-stage training framework to train the model.
Outcome: The proposed model achieves the state-of-the-art (SOTA) performance on four out of five language pairs.
A Massively Multilingual Analysis of Cross-linguality in Shared Embedding Space (2021.emnlp-main)

Copied to clipboard

Challenge: Cross-lingual language models house representations for many different languages in the same space.
Approach: They investigate linguistic and non-linguistic factors affecting sentence-level alignment in cross-lingual pretrained language models for 101 languages and 5,050 language pairs.
Outcome: The results show that word order agreement and agreement in morphological complexity are strongest predictors of cross-linguality.
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models.
Approach: They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key .
Outcome: The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key .
Aligning Cross-lingual Sentence Representations with Dual Momentum Contrast (2021.emnlp-main)

Copied to clipboard

Challenge: Existing work uses sentences within the same batch as negatives, which suffers from easy negatives.
Approach: They propose to align sentence representations from different languages into a unified embedding space . they adapt MoCo to further improve the quality of alignment .
Outcome: The proposed model achieves state-of-the-art on several tasks.
SentAlign: Accurate and Scalable Sentence Alignment (2023.emnlp-demo)

Copied to clipboard

Challenge: SentAlign is an automatic sentence alignment tool designed for large documents . it evaluates all possible alignment paths in documents of thousands of sentences .
Approach: They present a sentence alignment tool that evaluates all possible alignment paths in parallel documents of thousands of sentences and uses a divide-and-conquer approach to align documents containing tens of thousands.
Outcome: The proposed tool outperforms five other sentence alignment tools on two evaluation sets and on a downstream machine translation task.
SpanAlign: Sentence Alignment Method based on Cross-Language Span Prediction and ILP (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for automatic sentence alignment assume monotonic alignments, but they can handle non-monotonic alignments.
Approach: They propose a method to automatically extract parallel sentences from noisy parallel documents by embeddings and encoding each source and target sentence.
Outcome: The proposed method improves translation accuracy by 4.1 BLEU scores on English-Japanese . it can predict spans in target document from sentences in source document .
Improving Low-Resource Cross-lingual Document Retrieval by Reranking with Deep Bilingual Representations (P19-1)

Copied to clipboard

Challenge: Experimental results show that our model outperforms competitive translation-based baselines on cross-lingual relevance ranking tasks.
Approach: They propose to match queries and documents in both source and target languages with deep bilingual query-document representations.
Outcome: The proposed model outperforms translation-based baselines on English-Swahili, English-Tagalog, and English-Somali cross-lingual retrieval tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations