Human Raters Cannot Distinguish English Translations from Original English Texts (2023.emnlp-main)
Copied to clipboard
| Challenge: | Prior work on translationese has identified common hallmarks of translationeses, but human accuracy of identifying translated text is understudied. |
| Approach: | They perform an evaluation of English original/translated texts to examine whether raters can classify texts as being original or translated English and the features that lead rater to judge text as being translated. |
| Outcome: | The results provide critical insight into work in translation studies and context for assessments of translationese classifiers. |
Similar Papers
Lost in Translation, and Found: Detecting and Interpreting Translation Effects (2026.acl-long)
Copied to clipboard
Shira Wein, Anna Serbina, Jiyuan Ji, Nathan Wolf, Jason DeGraaff, Prajakta Kini, Maria Leonor Pacheco
| Challenge: | Translationese refers to the statistical patterns that distinguish translated texts from original texts. |
| Approach: | They analyze linguistic features which enable our model to achieve high accuracy by a collection of linguistic characteristics and pretrained neural models pick up these features without any fine-tuning. |
| Outcome: | The proposed model achieves high accuracy with a set of linguistic features that correspond to translationese theories and pretrained neural models pick up these features without any fine-tuning. |
Refined Assessment for Translation Evaluation: Rethinking Machine Translation Evaluation in the Era of Human-Level Systems (2025.findings-emnlp)
Copied to clipboard
Dmitry Popov, Vladislav Negodin, Ekaterina Enikeeva, Iana Matrosova, Nikolay Karpachev, Max Ryabinin
| Challenge: | Currently, traditional evaluation methods struggle to detect subtle translation errors. |
| Approach: | They propose to use a dataset of human evaluations for English–Russian translations created by professional linguists to enable consistent and rich annotation. |
| Outcome: | The proposed protocol allows expert assessments without time pressure to yield substantially different results from standard evaluations. |
Are we Estimating or Guesstimating Translation Quality? (2020.acl-main)
Copied to clipboard
| Challenge: | A carefully engineered ensemble of pre-trained multilingual language models won the QE shared task at WMT19. |
| Approach: | They propose to use pre-trained multilingual language models to train quality estimation for machine translation. |
| Outcome: | A carefully engineered ensemble of pre-trained language models wins the QE shared task at WMT19. |
Statistical Power and Translationese in Machine Translation Evaluation (2020.emnlp-main)
Copied to clipboard
| Challenge: | a recent paper argues that translationese has been used to describe features of translated text . a translationed text can be more explicit than the original source, authors say . authors recommend reverse-created test data be omitted from future evaluations . |
| Approach: | They propose to omit translationese from future machine translation evaluations . they also re-evaluate a past evaluation claiming human-parity of MT . |
| Outcome: | The proposed analysis shows that translationese does not affect machine translation evaluations. |
Train, Sort, Explain: Learning to Diagnose Translation Models (N19-4)
Copied to clipboard
| Challenge: | Evaluating translation models is a trade-off between effort and detail. |
| Approach: | They propose to use a neural text classifier to automatically expose systematic differences between human and machine translations to human experts. |
| Outcome: | The proposed method exposes systematic differences between human and machine translations to human experts. |
Can Automatic Metrics Assess High-Quality Translations? (2024.emnlp-main)
Copied to clipboard
| Challenge: | a recent human evaluation study found that translations produced by current MT systems achieve very high-quality scores when judged by humans on a direct assessment scale of 0 to 100. |
| Approach: | They stress-test the ability of current translation quality metrics to detect correct translations . they show that current metrics often over or underestimate translation quality . |
| Outcome: | The proposed method overestimates translation quality, the authors show . they show that current metrics often overestimate translation quality . |
Assessing Human-Parity in Machine Translation on the Segment Level (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent machine translation shared tasks have shown top-performing systems to tie or outperform human translation. |
| Approach: | They examine the outputs of top-performing systems in a recent machine translation shared task . they find that some systems outperform human translation on average . |
| Outcome: | a new method identifies segments for which human and machine perform poorly . the results show that top-performing systems outperform human translation on average . |
Has Machine Translation Achieved Human Parity? A Case for Document-level Evaluation (D18-1)
Copied to clipboard
| Challenge: | Recent research suggests that neural machine translation achieves parity with professional human translation on the WMT Chinese–English news translation task. |
| Approach: | They empirically test neural machine translation on a Chinese–English news translation task . they show human raters prefer human over machine translation when evaluating documents . |
| Outcome: | The proposed method shows that human translators prefer document-level evaluation over machine translation . the results highlight the need to shift towards document- level evaluation as machine translation improves . |
A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents (L18-1)
Copied to clipboard
| Challenge: | Terms are notoriously difficult to identify, both automatically and manually. |
| Approach: | They propose a method to annotate terms manually from a comparable corpus . they show that the gold standard provides a tool for evaluation and a rich source of information . |
| Outcome: | The proposed method provides a tool for evaluation and rich source of information about terms. |
Translationese as a Language in “Multilingual” NMT (2020.acl-main)
Copied to clipboard
| Challenge: | Recent work examines the impact of translationese in machine translation evaluation using the WMT evaluation campaign. |
| Approach: | They propose to use a sentence-level classifier to distinguish translationese from original target text to generate a machine translation model that can produce more natural outputs at test time. |
| Outcome: | The proposed model produces more natural outputs at test time, yielding gains in human evaluation scores on accuracy and fluency. |