| Challenge: | Annotation metrics are misaligned with the ideal measure of text quality and human evaluation remains the most accurate, reliable, and ultimate standard. |
| Approach: | They propose an annotation protocol that helps annotators mark erroneous parts of the translation and assign a final score. |
| Outcome: | The proposed protocol reduces the time per span annotation by half . the method reduces annotation budget by 25% with filtering of examples that the AI deems to be likely to be correct. |
Similar Papers
Refined Assessment for Translation Evaluation: Rethinking Machine Translation Evaluation in the Era of Human-Level Systems (2025.findings-emnlp)
Copied to clipboard
Dmitry Popov, Vladislav Negodin, Ekaterina Enikeeva, Iana Matrosova, Nikolay Karpachev, Max Ryabinin
| Challenge: | Currently, traditional evaluation methods struggle to detect subtle translation errors. |
| Approach: | They propose to use a dataset of human evaluations for English–Russian translations created by professional linguists to enable consistent and rich annotation. |
| Outcome: | The proposed protocol allows expert assessments without time pressure to yield substantially different results from standard evaluations. |
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress (2025.acl-short)
Copied to clipboard
| Challenge: | In machine translation evaluation, metric performance is assessed based on agreement with human judgments. |
| Approach: | They incorporate human baselines into the MT meta-evaluation to gain a clearer understanding of metric performance and establish an upper bound. |
| Outcome: | The results suggest human parity, but there are several reasons to caution . |
Computer Assisted Translation with Neural Quality Estimation and Automatic Post-Editing (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Using neural machine translation to approximate human parity is difficult due to the lack of parallel training corpora. |
| Approach: | They propose an end-to-end deep learning framework for quality estimation and automatic post-editing of machine translation output. |
| Outcome: | The proposed framework achieves state-of-the-art performance on the English–German dataset and human translators can significantly expedite their post-editing processing with the model. |
MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators (2025.coling-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown significant potential as judges for Machine Translation (MT) quality assessment. |
| Approach: | They propose a framework that automatically post-edits the original translation based on each error, thereby filtering out non-impactful errors. |
| Outcome: | The proposed framework improves reliability and quality of error spans against GEMBA-MQM, across eight LLMs in both high- and low-resource languages. |
Enhancing Human Evaluation in Machine Translation with Comparative Judgement (2025.acl-long)
Copied to clipboard
| Challenge: | Human evaluation is crucial for assessing rapidly evolving language models but is influenced by annotator proficiency and task design. |
| Approach: | They evaluate three annotation setups to integrate comparative judgment into human annotation for machine translation. |
| Outcome: | The proposed approach improves inter-annotator agreement and stability of the annotations. |
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? (2025.acl-long)
Copied to clipboard
| Challenge: | Pairwise feedback is widely used to evaluate and provide feedback to large language models (LLMs). |
| Approach: | They propose a tool-using agentic system to provide higher quality feedback on three challenging response domains: long-form factual, math and code tasks. |
| Outcome: | The proposed system can provide higher quality pairwise comparisons on three domains, independent of the LLM’s internal knowledge and biases. |
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation (2026.acl-long)
Copied to clipboard
| Challenge: | a critical component of machine translation model development is evaluating model quality. |
| Approach: | They propose a two-stage version of the current translation evaluation paradigm (MQM) they propose re-annotation, which uses raters to review and edit annotations . |
| Outcome: | The proposed method improves annotation quality by finding errors missed in the first pass. |
Automatic Correction of Human Translations (2022.naacl-main)
Copied to clipboard
| Challenge: | Despite recent advances in machine translation, a tremendous amount of translated content in the world is still written by humans. |
| Approach: | They propose a task of translation error correction (TEC) that corrects human-generated translations by correcting all errors in a source sentence and a human-created translation. |
| Outcome: | The proposed system improves translation accuracy by 5.1 points compared to MT systems with human errors . |
Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation (2021.tacl-1)
Copied to clipboard
| Challenge: | a large study of machine translation systems shows poor evaluation procedures can lead to erroneous conclusions. |
| Approach: | They propose an evaluation methodology grounded in explicit error analysis based on the Multidimensional Quality Metrics framework. |
| Outcome: | The proposed evaluation methodology outperforms crowd workers in two languages . it shows that human-based metrics outperformed crowd workers . |
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement (2025.emnlp-main)
Copied to clipboard
| Challenge: | Modern WQE techniques rely on expensive inference with large language models or ad-hoc training with large amounts of human-labeled data. |
| Approach: | They propose to use word-level quality estimation to identify translation errors from the inner workings of translation models to quantify the impact of human label variation on metric performance. |
| Outcome: | The proposed methods identify translation errors from the inner workings of translation models using human labels. |