Papers with English-Russian

12 papers
HOPE: A Task-Oriented and Human-Centric Evaluation Framework Using Professional Post-Editing Towards More Effective MT Evaluation (2022.lrec-1)

Copied to clipboard

Challenge: Existing automated evaluation metrics for machine translation are expensive and lack inter-rater reliability.
Approach: They propose a task-oriented and human-centric evaluation framework for machine translation output based on professional post-e diting annotations.
Outcome: The proposed framework improves translation quality and system performance and transparency . it is cost-effective, easy to use and faster to implement .
TransLLaMa: LLM-based Simultaneous Translation System (2024.findings-emnlp)

Copied to clipboard

Challenge: Decoder-only large language models have limited applications in simultaneous machine translation . naively translating each source word immediately results in compromised target quality .
Approach: a study shows that a pre-trained open-source LLM can control input segmentation directly by generating a special "wait" token.
Outcome: a new open-source model can control input segmentation directly by generating a special "wait" token.
Context-Aware Monolingual Repair for Neural Machine Translation (D19-1)

Copied to clipboard

Challenge: et al., 2018) show that human raters prefer corrected translations over the baseline ones.
Approach: They propose a monolingual model to correct inconsistencies between sentences . they use monolingual document-level data to train the model .
Outcome: The proposed model improves translations of contextual phenomena in English-Russian translation task.
When a Good Translation is Wrong in Context: Context-Aware Machine Translation Improves on Deixis, Ellipsis, and Lexical Cohesion (P19-1)

Copied to clipboard

Challenge: et al., 2018: translation errors due to the lack of extra-sentential context are becoming more and more noticeable among otherwise adequate translations.
Approach: They propose a context-aware translation model that uses sentence-level data to identify inconsistencies . standard metrics are not sensitive to improvements in consistency in document-level translations .
Outcome: The proposed model shows major gains over baseline without sacrificing performance . standard metrics are not sensitive to improvements in document-level translations .
Context-Aware Neural Machine Translation Learns Anaphora Resolution (P18-1)

Copied to clipboard

Challenge: Standard machine translation systems process sentences in isolation and ignore extra-sentential information.
Approach: They propose a context-aware neural machine translation model that controls flow of information from extended context to the translation model.
Outcome: The proposed model improves on an English-Russian subtitles dataset over its context-agnostic version (+0.7) and over simple concatenation of context and source sentences (+0.6).
Extract and Edit: An Alternative to Back-Translation for Unsupervised Neural Machine Translation (N19-1)

Copied to clipboard

Challenge: Back-translation has been used in previous approaches for unsupervised neural machine translation, but pseudo sentences are of low quality as translation errors accumulate during training.
Approach: They propose an approach to extract and edit real sentences from monolingual corpora and introduce a comparative translation loss to evaluate the translated target sentences.
Outcome: The proposed approach outperforms state-of-the-art translation systems across two benchmarks and two low-resource language pairs by more than 2 BLEU points.
DEEP: DEnoising Entity Pre-training for Neural Machine Translation (2022.acl-long)

Copied to clipboard

Challenge: Earlier named entity translation methods focus on phonetic transliteration, which ignores the sentence context for translation.
Approach: They propose a DEnoising Entity Pre-training method that leverages monolingual data and a knowledge base to improve named entity translation accuracy within sentences.
Outcome: The proposed method improves on three language pairs and denoising auto-encoding baselines.
On the Complementarity between Pre-Training and Back-Translation for Neural Machine Translation (2021.findings-emnlp)

Copied to clipboard

Challenge: Experimental results show that PT and BT are nicely complementary to each other.
Approach: They introduce two probing tasks for PT and BT respectively and investigate their complementarity.
Outcome: The proposed methods establish state-of-the-art on the WMT16 English-Romanian and English-Russian benchmarks.
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations (2024.findings-naacl)

Copied to clipboard

Challenge: supervised systems have not replaced dedicated supervised models for machine translation tasks.
Approach: They propose to guide LLMs to post-edit MT with feedback from MQM annotations . they then fine-tune the LLM to improve its ability to exploit the feedback .
Outcome: The proposed model improves TER, BLEU and COMET scores on Chinese-English, English-German and English-Russian data.
Refer to the Reference: Reference-focused Synthetic Automatic Post-Editing Data Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing approaches to synthetic APE data generation use source (src) sentences in a parallel corpus to obtain translations (mt) through an MT system and treat corresponding reference (ref) sentences as post-edits (pe).
Approach: They propose a reference-focused synthetic APE data generation technique that uses ‘ref’ instead of src’ sentences to obtain corrupted translations.
Outcome: The proposed technique improves on English-German, English-Russian, English -Marathi, English and Hindi language pairs.
Combining Word Embeddings with Bilingual Orthography Embeddings for Bilingual Dictionary Induction (2020.coling-main)

Copied to clipboard

Challenge: Bilingual dictionary induction (BDI) is a task of finding target language translations of source language words.
Approach: They propose to use bilingual orthography Embeddings to enrich BWE-based BDI with transliteration information to make a decision on which information source is more reliable for a particular word pair.
Outcome: The proposed system improves on English-Russian BDI and shows that it can be built with only weak bilingual signals and even without any bilingual signal.
Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned (P19-1)

Copied to clipboard

Challenge: et al., 2017) show that multi-head attention is important for neural machine translation.
Approach: They evaluate the contribution made by individual attention heads to the overall performance of the Transformer model and analyze the roles played by them in the encoder.
Outcome: The proposed pruning method removes the vast majority of heads without affecting performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations