Challenge: Existing studies have relied on out-of-the-box machine translation metrics to evaluate interpretation data, but they do not account for human judgments of interpretation quality.
Approach: They propose to use machine translation metrics to evaluate human interpretations to address potential barriers to disfluency, summarization, paraphrasing and segmentation.
Outcome: The proposed model achieves better correlation with human judgments than state-of-the-art metrics.

Similar Papers

Toward Machine Interpreting: Lessons from Human Interpreting Studies (2025.emnlp-main)

Copied to clipboard

Challenge: Current speech translation systems are static and do not adapt to real-world situations in ways human interpreters do.
Approach: They propose to model human interpreting using a new language model to improve usability . they argue that there is great potential to adopt many human interpreted principles .
Outcome: The proposed models can be used to improve human interpreting and improve translation performance.
Automatic Estimation of Simultaneous Interpreter Performance (P18-2)

Copied to clipboard

Challenge: Existing methods to predict interpreter confidence and the adequacy of the interpreted message are lacking.
Approach: They propose to extend a QE pipeline to estimate interpreter performance by using five settings in three language pairs.
Outcome: The proposed method can predict interpreter confidence and adequacy over five settings in three language pairs and improves interpretation strategy and evaluation measures.
Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for judging metrics are sensitive to the translations used for evaluation, leading to falsely confident conclusions about a metric’s efficacy.
Approach: They propose a method for thresholding performance improvement under an automatic metric against human judgements by using a pairwise system ranking method.
Outcome: The proposed method allows quantification of type I versus type II errors incurred, i.e., insignificant human differences in system quality that are accepted, and significant human differences that are rejected.
Refined Assessment for Translation Evaluation: Rethinking Machine Translation Evaluation in the Era of Human-Level Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: Currently, traditional evaluation methods struggle to detect subtle translation errors.
Approach: They propose to use a dataset of human evaluations for English–Russian translations created by professional linguists to enable consistent and rich annotation.
Outcome: The proposed protocol allows expert assessments without time pressure to yield substantially different results from standard evaluations.
Simultaneous Translation (2020.emnlp-tutorials)

Copied to clipboard

Challenge: Simultaneous translation is a problem that has long been considered one of the hardest problems in AI . this tutorial will provide a deep understanding of the history and the recent advances in simultaneous translation.
Approach: This tutorial will examine the design and evaluation of policies for simultaneous translation . it will provide an overview of the history and recent advances in simultaneous translation.
Outcome: This tutorial will examine the design and evaluation of policies for simultaneous translation .
Assessing Human-Parity in Machine Translation on the Segment Level (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent machine translation shared tasks have shown top-performing systems to tie or outperform human translation.
Approach: They examine the outputs of top-performing systems in a recent machine translation shared task . they find that some systems outperform human translation on average .
Outcome: a new method identifies segments for which human and machine perform poorly . the results show that top-performing systems outperform human translation on average .
Opportunities for Human-centered Evaluation of Machine Translation Systems (2022.findings-naacl)

Copied to clipboard

Challenge: a new study examines the role of machine translation in larger user-facing systems . a sysadmin and a human factors researcher are developing evaluation tools .
Approach: They argue that machine translation models are embedded in larger user-facing systems . they argue that evaluation at the systems level is still lacking .
Outcome: The proposed model evaluations are based on human-computer interaction models . the authors argue that evaluations should be based more on the entire system .
It Is Not As Good As You Think! Evaluating Simultaneous Machine Translation on Interpretation Data (2021.emnlp-main)

Copied to clipboard

Challenge: Existing siMT systems are trained and evaluated on offline translations . however, evaluation gap remains notable, calling for constructing large-scale interpretation corpora .
Approach: They propose a translation-to-interpretation transfer method which converts offline translations into interpretation-style data.
Outcome: The proposed interpretation test set shows that SiMT models improve on translation vs interpretation data.
DialSummEval: Revisiting Summarization Evaluation for Dialogues (2022.naacl-main)

Copied to clipboard

Challenge: Current models for dialogue summarization have flaws that may not be well exposed by frequently used metrics such as ROUGE.
Approach: They propose to re-evaluate 18 categories of metrics in terms of four dimensions: coherence, consistency, fluency and relevance, as well as a unified human evaluation of various models for the first time.
Outcome: The proposed dataset will be used to evaluate 18 categories of metrics in terms of coherence, consistency, fluency and relevance, and a unified human evaluation of various models for the first time.
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress (2025.acl-short)

Copied to clipboard

Challenge: In machine translation evaluation, metric performance is assessed based on agreement with human judgments.
Approach: They incorporate human baselines into the MT meta-evaluation to gain a clearer understanding of metric performance and establish an upper bound.
Outcome: The results suggest human parity, but there are several reasons to caution .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations