Challenge: Specific-domain bilingual lexicons are composed of MultiWord Expressions (MWEs) the manual construction of MWEs bilingual dictionaries is costly and time-consuming.
Approach: They propose to use word alignment approaches to automatically construct bilingual lexicons of MWEs from parallel corpora by formalizing the alignment process as an integer linear programming problem.
Outcome: The proposed approach extracts and aligns multiword expressions from parallel corpora and then filters them using linguistic patterns to build bilingual lexicons.

Similar Papers

Building Comparable Corpora for Assessing Multi-Word Term Alignment (2022.lrec-1)

Copied to clipboard

Challenge: Existing methods to extract bilingual terminologies from corpora are limited . MWTs pose serious challenges for alignment and machine translation systems .
Approach: They propose an approach to build comparable corpora and bilingual term dictionaries that evaluate bilingual term alignment in comparable corpus.
Outcome: The proposed method is validated on an existing dataset and manually annotated data.
MultiMWE: Building a Multi-lingual Multi-Word Expression (MWE) Parallel Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Existing bilingual or multi-lingual MWE corpora are limited for multilingual use . only 871 pairs of English-German MWEs are available for research .
Approach: They present a collection of bilingual and multi-lingual MWEs extracted from parallel corpora.
Outcome: The available bilingual or multi-lingual MWE corpus is very limited . the collection is a small collection of 871 pairs of English-German MWEs .
Towards a unified framework for bilingual terminology extraction of single-word and multi-word terms (C18-1)

Copied to clipboard

Challenge: Existing methods for extracting bilingual terminology from comparable corpora are limited to a set of syntactic patterns.
Approach: They propose a framework for aligning bilingual terms independently of term lengths . they introduce some enhancements to the context-based and neural network based approaches .
Outcome: The proposed framework improves the performance of the context-based and neural network based approaches and can be adapted in specialized domains.
Word Alignment by Fine-tuning Embeddings on Parallel Corpora (2021.eacl-main)

Copied to clipboard

Challenge: Existing work on word alignment has focused on unsupervised learning on parallel text.
Approach: They propose to combine pre-trained contextualized word embeddings with multilingually trained language models to achieve competitive results on word alignment tasks.
Outcome: The proposed model outperforms state-of-the-art models on five language pairs and can train multilingual word aligners that can obtain robust performance on different language pairs.
Bilingual Lexicon Induction through Unsupervised Machine Translation (P19-1)

Copied to clipboard

Challenge: Existing methods for bilingual lexicon induction use nearest neighbor or related retrieval methods to induce word translation pairs.
Approach: They propose a method that aligns word embeddings in two languages and uses them to build a phrase-table and a language model to extract the bilingual lexicon.
Outcome: The proposed method improves accuracy 6 points over nearest neighbor and 4 points over CSLS retrieval on the same cross-lingual embeddings.
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Recent studies have highlighted the potential of exploiting parallel corpora to enhance multilingual large language models.
Approach: They investigate the impact of parallel corpora quality and quantity, training objectives, and model size on performance of multilingual large language models enhanced with parallel corporeal.
Outcome: The proposed approach improves performance in bilingual and general-purpose tasks.
Pre-tokenization of Multi-word Expressions in Cross-lingual Word Embeddings (2020.emnlp-main)

Copied to clipboard

Challenge: Multi-Word Expressions (MWEs) are common in every language, but they are not translated by cross-lingual word embeddings.
Approach: They propose a method for word translation of Multi-Word Expressions (MWEs) they compile lists of MWEs in each language and tokenize them as single tokens before training word embeddings.
Outcome: The proposed method can translate multi-word expressions to and from English in 10 languages.
word2word: A Collection of Bilingual Lexicons for 3,564 Language Pairs (2020.lrec-1)

Copied to clipboard

Challenge: Our dataset provides top-k word translations in 3,564 (directed) language pairs across 62 languages in OpenSubtitles2018.
Approach: They propose a dataset and an open-source Python package for cross-lingual word translations extracted from sentence-level parallel corpora.
Outcome: The proposed bilingual lexicons have high coverage and achieve competitive translation quality for several language pairs.
Effectively Aligning and Filtering Parallel Corpora under Sparse Data Conditions (2020.acl-srw)

Copied to clipboard

Challenge: Parallel corpora are key to developing good machine translation systems, but abundant parallel data is hard to come by for languages with a low number of speakers.
Approach: They propose an unsupervised alignment method that can handle rich morphology by removing incorrect translations and segments containing extraneous data.
Outcome: The proposed method maximizes the number of correctly translated segments in a corpus and minimises noise by removing incorrect translations and segments containing extraneous data.
CoAM: Corpus of All-Type Multiword Expressions (2025.acl-long)

Copied to clipboard

Challenge: Existing datasets for multiword expressions are inconsistently annotated, limited to a single type of MWE, or limited in size.
Approach: They propose to use a new interface to generate MWE annotations for the first time in a dataset of MWE identification.
Outcome: The proposed model outperforms existing models on the DiMSUM dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations