Challenge: Recent research has focused on literary machine translation (MT) but evaluation of literary MT remains an open problem.
Approach: They propose a paragraph-level parallel corpus containing verified human translations and 13k evaluated sentences across four language pairs.
Outcome: The proposed corpus compares human evaluations with students and professionals . it shows that the adequacy of human evaluation is controlled by two factors .

Similar Papers

LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question Answering (2025.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics for literature prioritize mechanical accuracy over artistic expression . this bias could result in an irreversible decline in translation quality and cultural authenticity .
Approach: They propose a novel, reference-free, LLM-based question-answering framework for literary translation evaluation.
Outcome: a novel, reference-free, LLM-based question-answering framework is developed for literary translation evaluation.
Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World Literature (2022.emnlp-main)

Copied to clipboard

Challenge: Literary translation is a culturally significant task, but it is bottlenecked by the small number of qualified literary translators . a dataset of non-English language novels is used to study literary MT .
Approach: They use a dataset of non-English language novels aligned to human and automatic English translations to study literary MT.
Outcome: The proposed model prefers human translations over machine translations at a rate of 84% . state-of-the-art MT metrics do not correlate with preferences, the study finds .
Refined Assessment for Translation Evaluation: Rethinking Machine Translation Evaluation in the Era of Human-Level Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: Currently, traditional evaluation methods struggle to detect subtle translation errors.
Approach: They propose to use a dataset of human evaluations for English–Russian translations created by professional linguists to enable consistent and rich annotation.
Outcome: The proposed protocol allows expert assessments without time pressure to yield substantially different results from standard evaluations.
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress (2025.acl-short)

Copied to clipboard

Challenge: In machine translation evaluation, metric performance is assessed based on agreement with human judgments.
Approach: They incorporate human baselines into the MT meta-evaluation to gain a clearer understanding of metric performance and establish an upper bound.
Outcome: The results suggest human parity, but there are several reasons to caution .
What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary Translation (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have made document-level machine translation increasingly practical, enabled by long-context modeling and strong generation quality.
Approach: They propose to use document-level MT followed by segment-level refinement to find the strongest and most stable improvements across six LLMs and seven language pairs.
Outcome: The proposed method outperforms error-specific prompting and evaluate-then-refine schemes in document-level translation.
Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have advanced machine translation (MT) a meta-evaluation dataset focused on non-literal translations is lacking . experimental results show the inaccuracies of traditional MT metrics and the limitations of LLM-as-a-Judge.
Approach: They propose a meta-evaluation framework that leverages sub-agents to evaluate machine translation metrics.
Outcome: The proposed framework improves on the knowledge cutoff and score inconsistency problem.
Towards A “Novel” Benchmark: Evaluating Literary Fiction with Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) context windows have enabled them to process inputs over 100K tokens and generate outputs of up to 10K token.
Approach: They propose a multi-level evaluation framework that incorporates ten metrics across the Macro, Meso, and Micro levels and an annotated fiction dataset.
Outcome: The proposed framework incorporates ten metrics across the Macro, Meso, and Micro levels and is based on a human-human-AI dataset.
Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation (2021.tacl-1)

Copied to clipboard

Challenge: a large study of machine translation systems shows poor evaluation procedures can lead to erroneous conclusions.
Approach: They propose an evaluation methodology grounded in explicit error analysis based on the Multidimensional Quality Metrics framework.
Outcome: The proposed evaluation methodology outperforms crowd workers in two languages . it shows that human-based metrics outperformed crowd workers .
Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for judging metrics are sensitive to the translations used for evaluation, leading to falsely confident conclusions about a metric’s efficacy.
Approach: They propose a method for thresholding performance improvement under an automatic metric against human judgements by using a pairwise system ranking method.
Outcome: The proposed method allows quantification of type I versus type II errors incurred, i.e., insignificant human differences in system quality that are accepted, and significant human differences that are rejected.
Benchmarking and Improving Long-Text Translation with Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies have illuminated the promising capabilities of large language models (LLMs) in handling long texts.
Approach: They construct a benchmark dataset specifically designed for the finetuning and evaluation of large language models (LLMs) they compare LLMs with MT models and find they exhibit shortcomings in long-text domains .
Outcome: The proposed model performs better in long-text translation, and its performance diminishes as document size increases.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations