Challenge: Current approaches to incorporating terminology constraints in machine translation (MT) typically assume that the constraint terms are provided in their correct morphological forms.
Approach: They propose a framework for incorporating lemma constraints in machine translation . they use a cross-lingual inflection module that inflects the target lemmo constraints based on the source context.
Outcome: The proposed framework outperforms existing methods with lower training costs and linguistic knowledge in domain adaptation and low-resource MT settings.

Similar Papers

End-to-End Lexically Constrained Machine Translation for Morphologically Rich Languages (2021.acl-long)

Copied to clipboard

Challenge: Existing approaches to enforce word forms in translations struggle to make them agree with the rest of the output.
Approach: They propose to train neural machine translation models with lemmatized constraints to infer correct word inflection.
Outcome: The proposed model reduces errors in translation of constrained terms in automatic and manual evaluations on English-Czech language pairs.
Morphology Aware Source Term Masking for Terminology-Constrained NMT (2024.findings-eacl)

Copied to clipboard

Challenge: Recent research in terminology-constrained NMT systems focuses on data-driven approaches to generating translations.
Approach: They propose a method that appends target term lemmas to their corresponding source terms in the input sentence while retaining essential grammatical information.
Outcome: The proposed method improves on the “copy-and-inflect” method in two translation directions with different levels of source morphological complexity.
Training Neural Machine Translation to Apply Terminology Constraints (P19-1)

Copied to clipboard

Challenge: Existing methods to integrate domain terminology into neural machine translation (NMT) are brittle when tested in real-world situations.
Approach: They propose a method to inject custom terminology into neural machine translation at run time by using the target side of terminology entries whose source side match the input as decoding-time constraints.
Outcome: The proposed method is faster than state-of-the-art decoding and more efficient than constraint-free decoding.
Alignment verification to improve NMT translation towards highly inflectional languages with limited resources (2021.eacl-main)

Copied to clipboard

Challenge: Existing approaches to improve translation quality using limited training data are phrase-based and syntax-based approaches.
Approach: They propose to combine a neural MT system with an open source module to improve translation quality.
Outcome: The proposed method improves translation quality over the best individual NMT and the standard ensemble system provided in the Marian-NMT system.
Neural Machine Translation Decoding with Terminology Constraints (N18-2)

Copied to clipboard

Challenge: Constrained neural machine translation systems can provide excellent quality but do not strictly enforce terminology.
Approach: They propose a framework for constrained neural decoding which supports target-side constraints as well as constraints with corresponding aligned input text spans.
Outcome: The proposed framework performs well on multiple translation tasks and motivates the need for constrained decoding with attentions to reduce misplacement and duplication when translating user constraints.
Improving Character-Based Decoding Using Target-Side Morphological Information for Neural Machine Translation (N18-1)

Copied to clipboard

Challenge: Morphologically complex words (MCWs) are multi-layer structures consisting of different subunits, each of which carries semantic information and has a specific syntactic role.
Approach: They propose an extension to the state-of-the-art model which works at the character level and boosts the decoder with target-side morphological information.
Outcome: The proposed model improves on the state-of-the-art model and can be extended to include morphologically complex words (MCWs) in three languages.
A Template-based Method for Constrained Neural Machine Translation (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods to solve this problem can not satisfy the following three desiderata: (1) high translation quality, (2) high match accuracy, and (3) low latency.
Approach: They propose a template-based method that can provide high translation quality and match accuracy and a low latency inference.
Outcome: The proposed method outperforms baselines in lexically and structurally constrained translation tasks and can be used in a variety of applications.
Tailoring Neural Architectures for Translating from Morphologically Rich Languages (C18-1)

Copied to clipboard

Challenge: A morphologically complex word is a hierarchical constituent with meaning-preserving subunits, so word-based models which rely on surface forms might not be powerful enough to translate such structures.
Approach: They propose a neural architecture which is designed to deal with morphological complexities on the source side and redesign the decoder accordingly to benefit from such information.
Outcome: The proposed model outperforms existing subword- and character-based architectures and showed significant improvements on translating from German, Russian, and Turkish into English.
Evaluating Pre-training Objectives for Low-Resource Translation into Morphologically Rich Languages (2022.lrec-1)

Copied to clipboard

Challenge: a lack of parallel data is a major limitation for Neural Machine Translation systems, especially for morphologically rich languages.
Approach: They propose to leverage target monolingual data to overcome the lack of parallel data . they introduce a new technique called PT-Inflect to train NMT systems .
Outcome: The proposed techniques outperform NMT systems trained on parallel data on four typologically diverse target languages.
Integrating Vectorized Lexical Constraints for Neural Machine Translation (2022.acl-long)

Copied to clipboard

Challenge: Existing studies focus on integrating discrete lexical constraints into neural machine translation models.
Approach: They propose to integrate constraints into NMT models by integrating them into keys and values . they show that their method outperforms representative baselines on four language pairs .
Outcome: The proposed method outperforms baselines on four language pairs, showing superiority .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations