Challenge: Prior work on translationese has identified common hallmarks of translationeses, but human accuracy of identifying translated text is understudied.
Approach: They perform an evaluation of English original/translated texts to examine whether raters can classify texts as being original or translated English and the features that lead rater to judge text as being translated.
Outcome: The results provide critical insight into work in translation studies and context for assessments of translationese classifiers.

Similar Papers

Lost in Translation, and Found: Detecting and Interpreting Translation Effects (2026.acl-long)

Copied to clipboard

Challenge: Translationese refers to the statistical patterns that distinguish translated texts from original texts.
Approach: They analyze linguistic features which enable our model to achieve high accuracy by a collection of linguistic characteristics and pretrained neural models pick up these features without any fine-tuning.
Outcome: The proposed model achieves high accuracy with a set of linguistic features that correspond to translationese theories and pretrained neural models pick up these features without any fine-tuning.
Refined Assessment for Translation Evaluation: Rethinking Machine Translation Evaluation in the Era of Human-Level Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: Currently, traditional evaluation methods struggle to detect subtle translation errors.
Approach: They propose to use a dataset of human evaluations for English–Russian translations created by professional linguists to enable consistent and rich annotation.
Outcome: The proposed protocol allows expert assessments without time pressure to yield substantially different results from standard evaluations.
Are we Estimating or Guesstimating Translation Quality? (2020.acl-main)

Copied to clipboard

Challenge: A carefully engineered ensemble of pre-trained multilingual language models won the QE shared task at WMT19.
Approach: They propose to use pre-trained multilingual language models to train quality estimation for machine translation.
Outcome: A carefully engineered ensemble of pre-trained language models wins the QE shared task at WMT19.
Statistical Power and Translationese in Machine Translation Evaluation (2020.emnlp-main)

Copied to clipboard

Challenge: a recent paper argues that translationese has been used to describe features of translated text . a translationed text can be more explicit than the original source, authors say . authors recommend reverse-created test data be omitted from future evaluations .
Approach: They propose to omit translationese from future machine translation evaluations . they also re-evaluate a past evaluation claiming human-parity of MT .
Outcome: The proposed analysis shows that translationese does not affect machine translation evaluations.
Train, Sort, Explain: Learning to Diagnose Translation Models (N19-4)

Copied to clipboard

Challenge: Evaluating translation models is a trade-off between effort and detail.
Approach: They propose to use a neural text classifier to automatically expose systematic differences between human and machine translations to human experts.
Outcome: The proposed method exposes systematic differences between human and machine translations to human experts.
Can Automatic Metrics Assess High-Quality Translations? (2024.emnlp-main)

Copied to clipboard

Challenge: a recent human evaluation study found that translations produced by current MT systems achieve very high-quality scores when judged by humans on a direct assessment scale of 0 to 100.
Approach: They stress-test the ability of current translation quality metrics to detect correct translations . they show that current metrics often over or underestimate translation quality .
Outcome: The proposed method overestimates translation quality, the authors show . they show that current metrics often overestimate translation quality .
Assessing Human-Parity in Machine Translation on the Segment Level (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent machine translation shared tasks have shown top-performing systems to tie or outperform human translation.
Approach: They examine the outputs of top-performing systems in a recent machine translation shared task . they find that some systems outperform human translation on average .
Outcome: a new method identifies segments for which human and machine perform poorly . the results show that top-performing systems outperform human translation on average .
Has Machine Translation Achieved Human Parity? A Case for Document-level Evaluation (D18-1)

Copied to clipboard

Challenge: Recent research suggests that neural machine translation achieves parity with professional human translation on the WMT Chinese–English news translation task.
Approach: They empirically test neural machine translation on a Chinese–English news translation task . they show human raters prefer human over machine translation when evaluating documents .
Outcome: The proposed method shows that human translators prefer document-level evaluation over machine translation . the results highlight the need to shift towards document- level evaluation as machine translation improves .
A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents (L18-1)

Copied to clipboard

Challenge: Terms are notoriously difficult to identify, both automatically and manually.
Approach: They propose a method to annotate terms manually from a comparable corpus . they show that the gold standard provides a tool for evaluation and a rich source of information .
Outcome: The proposed method provides a tool for evaluation and rich source of information about terms.
Translationese as a Language in “Multilingual” NMT (2020.acl-main)

Copied to clipboard

Challenge: Recent work examines the impact of translationese in machine translation evaluation using the WMT evaluation campaign.
Approach: They propose to use a sentence-level classifier to distinguish translationese from original target text to generate a machine translation model that can produce more natural outputs at test time.
Outcome: The proposed model produces more natural outputs at test time, yielding gains in human evaluation scores on accuracy and fluency.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations