Challenge: Current machine translation techniques are bottlenecked by adequacy issues . we propose automatic detection of missing and wrong translations .
Approach: They propose automatic detection of adequacy errors in MT hypothesis for MT model evaluation by annotating missing and wrong translations in 15000 Chinese-English translation pairs.
Outcome: The proposed model can detect missing and wrong translations in 15000 Chinese-English translation pairs.

Similar Papers

Informative Manual Evaluation of Machine Translation Output (2020.coling-main)

Copied to clipboard

Challenge: a new method for manual evaluation of machine translation output is proposed . evaluators mark problematic parts of the translated text, not just overall scores .
Approach: They propose a method for manual evaluation of machine translation output based on marking actual issues in the translated text.
Outcome: The proposed method can be applied on any genre/domain and language pair . it can be guided by various types of quality criteria and can be used for other types of generated text.
An Alignment-Agnostic Model for Chinese Text Error Correction (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing models for Chinese text error correction can correct mistaken, missing and redundant characters, but they cannot handle missing or redundant characters.
Approach: They propose an alignment-agnostic framework to correct Chinese text errors . framework detects missing and redundant characters and can be used as a cold start model .
Outcome: The proposed framework can handle both text aligned and non-aligned situations and can serve as a cold start model when no annotation data are provided.
Toward Machine Translation Literacy: How Lay Users Perceive and Rely on Imperfect Translations (2025.emnlp-main)

Copied to clipboard

Challenge: Using machine translation tools for everyday tasks is becoming more commonplace, but a lack of evaluation strategies and alternatives can cause users to over-rely on it.
Approach: They propose to use MT evaluation techniques to promote MT quality and MT literacy among its users.
Outcome: The findings highlight the need for evaluation and NLP explanation techniques to promote MT quality and MT literacy among its users.
Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Machine-translated benchmarks are widely used to assess the multilingual capabilities of large language models (LLMs), yet translation errors in these benchmarks remain underexplored.
Approach: They show how well machine-translated benchmarks match human span annotations on translations . they also show how strongly translation errors explain accuracy drops on translated benchmarks - a gap that is not addressed yet .
Outcome: The proposed model matches human-level translations with human-language annotations on translations, but translation errors are associated with accuracy drops even after controlling for English correctness and source-side anomalies.
Evaluating Automatic Metrics with Incremental Machine Translation Systems (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have shown that neural metrics are more reliable than non-neural metrics.
Approach: They propose to use commercial machine translations to evaluate machine translation metrics based on their preference for more recent outputs.
Outcome: The proposed dataset confirms several previous findings, including the advantage of neural metrics over non-neural ones, and also explores the debated issue of how MT quality affects metric reliability.
Train, Sort, Explain: Learning to Diagnose Translation Models (N19-4)

Copied to clipboard

Challenge: Evaluating translation models is a trade-off between effort and detail.
Approach: They propose to use a neural text classifier to automatically expose systematic differences between human and machine translations to human experts.
Outcome: The proposed method exposes systematic differences between human and machine translations to human experts.
A fine-grained error analysis of NMT, SMT and RBMT output for English-to-Dutch (L18-1)

Copied to clipboard

Challenge: Since 2016, the landscape of automated translation has substantially changed with the arrival of neural machine translation (NMT).
Approach: They propose to use an annotated SCATE corpus of MT errors to enrich the SCATE error taxonomy to fit the neural MT output.
Outcome: The proposed system outperforms phrase-based and rule-based systems except for lexical issues.
Upping the Ante: Towards a Better Benchmark for Chinese-to-English Machine Translation (L18-1)

Copied to clipboard

Challenge: Currently, there is no widely accepted standard for evaluation of machine translation (MT) for Chinese-to-English translation, there are no standard for standardized training sets, development sets, and test sets.
Approach: They propose to use Chinese-to-English machine translation as a benchmark . they build a highly competitive state-of-the-art MT system that outperforms reported results .
Outcome: The proposed system outperforms reported results on NIST OpenMT test sets in almost all papers published in major conferences and journals in computational linguistics and artificial intelligence in the past 11 years.
Automatic Correction of Human Translations (2022.naacl-main)

Copied to clipboard

Challenge: Despite recent advances in machine translation, a tremendous amount of translated content in the world is still written by humans.
Approach: They propose a task of translation error correction (TEC) that corrects human-generated translations by correcting all errors in a source sentence and a human-created translation.
Outcome: The proposed system improves translation accuracy by 5.1 points compared to MT systems with human errors .
Lost in Translation, and Found: Detecting and Interpreting Translation Effects (2026.acl-long)

Copied to clipboard

Challenge: Translationese refers to the statistical patterns that distinguish translated texts from original texts.
Approach: They analyze linguistic features which enable our model to achieve high accuracy by a collection of linguistic characteristics and pretrained neural models pick up these features without any fine-tuning.
Outcome: The proposed model achieves high accuracy with a set of linguistic features that correspond to translationese theories and pretrained neural models pick up these features without any fine-tuning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations