HOPE: A Task-Oriented and Human-Centric Evaluation Framework Using Professional Post-Editing Towards More Effective MT Evaluation (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing automated evaluation metrics for machine translation are expensive and lack inter-rater reliability. |
| Approach: | They propose a task-oriented and human-centric evaluation framework for machine translation output based on professional post-e diting annotations. |
| Outcome: | The proposed framework improves translation quality and system performance and transparency . it is cost-effective, easy to use and faster to implement . |
Similar Papers
Refined Assessment for Translation Evaluation: Rethinking Machine Translation Evaluation in the Era of Human-Level Systems (2025.findings-emnlp)
Copied to clipboard
Dmitry Popov, Vladislav Negodin, Ekaterina Enikeeva, Iana Matrosova, Nikolay Karpachev, Max Ryabinin
| Challenge: | Currently, traditional evaluation methods struggle to detect subtle translation errors. |
| Approach: | They propose to use a dataset of human evaluations for English–Russian translations created by professional linguists to enable consistent and rich annotation. |
| Outcome: | The proposed protocol allows expert assessments without time pressure to yield substantially different results from standard evaluations. |
TRANSLATIONCORRECT: A Unified Framework for Machine Translation Post-Editing with Predictive Error Assistance (2025.acl-demo)
Copied to clipboard
| Challenge: | Current workflows for machine translation (MT) post-editing and research data collection are inefficient and time-consuming. |
| Approach: | They propose a framework that combines MT and error prediction within a single environment. |
| Outcome: | **TranslationCorrect** exports high-quality span-based annotations in the Error Span Annotation format, using an error taxonomy inspired by Multidimensional Quality Metrics (MQM). |
ReMedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling (2025.emnlp-main)
Copied to clipboard
| Challenge: | Regression-based neural metrics struggle with inconsistency of human ratings . prompting large language models (LLMs) for MT scoring has also shown promise . |
| Approach: | They propose a MT metric framework that reformulates translation evaluation as a reward modeling task. |
| Outcome: | The proposed framework surpasses larger WMT winners and massive closed LLMs across 39 language pairs and 111 MT systems. |
Rethinking the Word-level Quality Estimation for Machine Translation from Human Judgement (2023.findings-acl)
Copied to clipboard
| Challenge: | Word-level Quality Estimation (QE) of Machine Translation aims to detect potential translation errors in the translated sentence without reference. |
| Approach: | They propose to use a human-generated translation judgment to generate a word-level quality estimate (QE) using a translation error rate toolkit to detect translation errors without reference. |
| Outcome: | The proposed dataset is more consistent with human judgment and confirms the effectiveness of the proposed tag-correcting strategies. |
Extrinsic Evaluation of Machine Translation Metrics (2023.acl-long)
Copied to clipboard
| Challenge: | MT metrics are widely used to distinguish the quality of machine translation systems across relatively large test sets. |
| Approach: | They evaluate the segment-level performance of the most widely used MT metrics by correlating them with how useful they are for downstream tasks. |
| Outcome: | The MT metrics are widely used to distinguish the quality of machine translation systems across relatively large test sets. |
Enhancing Human Evaluation in Machine Translation with Comparative Judgement (2025.acl-long)
Copied to clipboard
| Challenge: | Human evaluation is crucial for assessing rapidly evolving language models but is influenced by annotator proficiency and task design. |
| Approach: | They evaluate three annotation setups to integrate comparative judgment into human annotation for machine translation. |
| Outcome: | The proposed approach improves inter-annotator agreement and stability of the annotations. |
Evaluating Automatic Metrics with Incremental Machine Translation Systems (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have shown that neural metrics are more reliable than non-neural metrics. |
| Approach: | They propose to use commercial machine translations to evaluate machine translation metrics based on their preference for more recent outputs. |
| Outcome: | The proposed dataset confirms several previous findings, including the advantage of neural metrics over non-neural ones, and also explores the debated issue of how MT quality affects metric reliability. |
Multi-Dimensional Machine Translation Evaluation: Model Evaluation and Resource for Korean (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies on MT evaluation characterize quality of output with a single number . a recent advancement in MT technologies has enabled higher-quality, more nuanced translations . |
| Approach: | They propose a 1200-sentence MQM evaluation benchmark for English-Korean and a reference-free QE setup to evaluate the quality of the translations. |
| Outcome: | The proposed model outperforms the existing model in style and accuracy. |
Measuring Uncertainty in Translation Quality Evaluation (TQE) (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing automated tools are not good enough to evaluate translation quality . existing tools are often accused of having low reliability and agreement . |
| Approach: | They propose to use a method to accurately estimate the confidence intervals depending on the sample size of the translated text. |
| Outcome: | The proposed method aims to estimate the confidence intervals (CITATION) depending on the sample size of the translated text, e.g. the amount of words or sentences, that needs to be processed on TQE workflow step for confident and reliable evaluation of overall translation quality. |
MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators (2025.coling-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown significant potential as judges for Machine Translation (MT) quality assessment. |
| Approach: | They propose a framework that automatically post-edits the original translation based on each error, thereby filtering out non-impactful errors. |
| Outcome: | The proposed framework improves reliability and quality of error spans against GEMBA-MQM, across eight LLMs in both high- and low-resource languages. |