Challenge: Annotation metrics are misaligned with the ideal measure of text quality and human evaluation remains the most accurate, reliable, and ultimate standard.
Approach: They propose an annotation protocol that helps annotators mark erroneous parts of the translation and assign a final score.
Outcome: The proposed protocol reduces the time per span annotation by half . the method reduces annotation budget by 25% with filtering of examples that the AI deems to be likely to be correct.

Similar Papers

Refined Assessment for Translation Evaluation: Rethinking Machine Translation Evaluation in the Era of Human-Level Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: Currently, traditional evaluation methods struggle to detect subtle translation errors.
Approach: They propose to use a dataset of human evaluations for English–Russian translations created by professional linguists to enable consistent and rich annotation.
Outcome: The proposed protocol allows expert assessments without time pressure to yield substantially different results from standard evaluations.
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress (2025.acl-short)

Copied to clipboard

Challenge: In machine translation evaluation, metric performance is assessed based on agreement with human judgments.
Approach: They incorporate human baselines into the MT meta-evaluation to gain a clearer understanding of metric performance and establish an upper bound.
Outcome: The results suggest human parity, but there are several reasons to caution .
Computer Assisted Translation with Neural Quality Estimation and Automatic Post-Editing (2020.findings-emnlp)

Copied to clipboard

Challenge: Using neural machine translation to approximate human parity is difficult due to the lack of parallel training corpora.
Approach: They propose an end-to-end deep learning framework for quality estimation and automatic post-editing of machine translation output.
Outcome: The proposed framework achieves state-of-the-art performance on the English–German dataset and human translators can significantly expedite their post-editing processing with the model.
MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown significant potential as judges for Machine Translation (MT) quality assessment.
Approach: They propose a framework that automatically post-edits the original translation based on each error, thereby filtering out non-impactful errors.
Outcome: The proposed framework improves reliability and quality of error spans against GEMBA-MQM, across eight LLMs in both high- and low-resource languages.
Enhancing Human Evaluation in Machine Translation with Comparative Judgement (2025.acl-long)

Copied to clipboard

Challenge: Human evaluation is crucial for assessing rapidly evolving language models but is influenced by annotator proficiency and task design.
Approach: They evaluate three annotation setups to integrate comparative judgment into human annotation for machine translation.
Outcome: The proposed approach improves inter-annotator agreement and stability of the annotations.
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? (2025.acl-long)

Copied to clipboard

Challenge: Pairwise feedback is widely used to evaluate and provide feedback to large language models (LLMs).
Approach: They propose a tool-using agentic system to provide higher quality feedback on three challenging response domains: long-form factual, math and code tasks.
Outcome: The proposed system can provide higher quality pairwise comparisons on three domains, independent of the LLM’s internal knowledge and biases.
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation (2026.acl-long)

Copied to clipboard

Challenge: a critical component of machine translation model development is evaluating model quality.
Approach: They propose a two-stage version of the current translation evaluation paradigm (MQM) they propose re-annotation, which uses raters to review and edit annotations .
Outcome: The proposed method improves annotation quality by finding errors missed in the first pass.
Automatic Correction of Human Translations (2022.naacl-main)

Copied to clipboard

Challenge: Despite recent advances in machine translation, a tremendous amount of translated content in the world is still written by humans.
Approach: They propose a task of translation error correction (TEC) that corrects human-generated translations by correcting all errors in a source sentence and a human-created translation.
Outcome: The proposed system improves translation accuracy by 5.1 points compared to MT systems with human errors .
Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation (2021.tacl-1)

Copied to clipboard

Challenge: a large study of machine translation systems shows poor evaluation procedures can lead to erroneous conclusions.
Approach: They propose an evaluation methodology grounded in explicit error analysis based on the Multidimensional Quality Metrics framework.
Outcome: The proposed evaluation methodology outperforms crowd workers in two languages . it shows that human-based metrics outperformed crowd workers .
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement (2025.emnlp-main)

Copied to clipboard

Challenge: Modern WQE techniques rely on expensive inference with large language models or ad-hoc training with large amounts of human-labeled data.
Approach: They propose to use word-level quality estimation to identify translation errors from the inner workings of translation models to quantify the impact of human label variation on metric performance.
Outcome: The proposed methods identify translation errors from the inner workings of translation models using human labels.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations