Challenge: In practice, LLMs are largely diagnostic, with the signals rarely translating into direct quality improvements under real production constraints.
Approach: They propose a two-stage, evaluator-guided automatic post-editing framework that turns MQM-style evaluation into targeted repairs.
Outcome: The proposed framework improves both COMET and CometKiwi scores over one-stage evaluation methods while severities and error spans show strong agreement with human annotations and human editor preferences.

Similar Papers

MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown significant potential as judges for Machine Translation (MT) quality assessment.
Approach: They propose a framework that automatically post-edits the original translation based on each error, thereby filtering out non-impactful errors.
Outcome: The proposed framework improves reliability and quality of error spans against GEMBA-MQM, across eight LLMs in both high- and low-resource languages.
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations (2024.findings-naacl)

Copied to clipboard

Challenge: supervised systems have not replaced dedicated supervised models for machine translation tasks.
Approach: They propose to guide LLMs to post-edit MT with feedback from MQM annotations . they then fine-tune the LLM to improve its ability to exploit the feedback .
Outcome: The proposed model improves TER, BLEU and COMET scores on Chinese-English, English-German and English-Russian data.
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation (2026.acl-long)

Copied to clipboard

Challenge: a critical component of machine translation model development is evaluating model quality.
Approach: They propose a two-stage version of the current translation evaluation paradigm (MQM) they propose re-annotation, which uses raters to review and edit annotations .
Outcome: The proposed method improves annotation quality by finding errors missed in the first pass.
RUBRIC-MQM : Span-Level LLM-as-judge in Machine Translation For High-End Models (2025.acl-industry)

Copied to clipboard

Challenge: Existing LLMs are unable to match outputs due to their open-ended nature .
Approach: They propose a meta-evaluation strategy PromptCUE to evaluate cutting-edge LAJ-MT models such as GEMBA-MQM and a rubric-style prompt tailored to the characteristics of LLMs.
Outcome: The proposed model is able to predict scores or identify errors for individual sentences and is reliable in the real world.
What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary Translation (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have made document-level machine translation increasingly practical, enabled by long-context modeling and strong generation quality.
Approach: They propose to use document-level MT followed by segment-level refinement to find the strongest and most stable improvements across six LLMs and seven language pairs.
Outcome: The proposed method outperforms error-specific prompting and evaluate-then-refine schemes in document-level translation.
TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved impressive results in Machine Translation (MT). human evaluations reveal that LLM-generated translations still contain various errors.
Approach: They propose a LLM-based self-refinement framework that feeds error information back into LLMs to facilitate self-finement, leading to enhanced translation quality.
Outcome: The proposed framework outperforms internal refinement and feedback methods while ensuring a robust translation quality baseline.
Contextual Refinement of Translations: Large Language Models for Sentence and Document-Level Post-Editing (2024.naacl-long)

Copied to clipboard

Challenge: Large language models have demonstrated considerable success in various natural language processing tasks, but their performance in NMT tasks is still underexplored.
Approach: They propose to use LLMs as automatic post-editors rather than direct translators to improve BLEU and COMET performance.
Outcome: The proposed approach improves BLEU but COMET performance compared to in-context learning.
Refined Assessment for Translation Evaluation: Rethinking Machine Translation Evaluation in the Era of Human-Level Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: Currently, traditional evaluation methods struggle to detect subtle translation errors.
Approach: They propose to use a dataset of human evaluations for English–Russian translations created by professional linguists to enable consistent and rich annotation.
Outcome: The proposed protocol allows expert assessments without time pressure to yield substantially different results from standard evaluations.
TRANSLATIONCORRECT: A Unified Framework for Machine Translation Post-Editing with Predictive Error Assistance (2025.acl-demo)

Copied to clipboard

Challenge: Current workflows for machine translation (MT) post-editing and research data collection are inefficient and time-consuming.
Approach: They propose a framework that combines MT and error prediction within a single environment.
Outcome: **TranslationCorrect** exports high-quality span-based annotations in the Error Span Annotation format, using an error taxonomy inspired by Multidimensional Quality Metrics (MQM).
Harnessing Large Language Models as Post-hoc Correctors (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have demonstrated their effectiveness in a wide range of tasks, including machine translation and commonsense reasoning.
Approach: They propose a training-free framework that can work as a post-hoc corrector to propose corrections for ML models.
Outcome: The proposed framework improves the performance of a number of models by up to 39% on text analysis and the challenging molecular predictions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations