Papers with GEMBA-MQM
RUBRIC-MQM : Span-Level LLM-as-judge in Machine Translation For High-End Models (2025.acl-industry)
Copied to clipboard
| Challenge: | Existing LLMs are unable to match outputs due to their open-ended nature . |
| Approach: | They propose a meta-evaluation strategy PromptCUE to evaluate cutting-edge LAJ-MT models such as GEMBA-MQM and a rubric-style prompt tailored to the characteristics of LLMs. |
| Outcome: | The proposed model is able to predict scores or identify errors for individual sentences and is reliable in the real world. |
MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators (2025.coling-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown significant potential as judges for Machine Translation (MT) quality assessment. |
| Approach: | They propose a framework that automatically post-edits the original translation based on each error, thereby filtering out non-impactful errors. |
| Outcome: | The proposed framework improves reliability and quality of error spans against GEMBA-MQM, across eight LLMs in both high- and low-resource languages. |