Barriers to Effective Evaluation of Simultaneous Interpretation (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing studies have relied on out-of-the-box machine translation metrics to evaluate interpretation data, but they do not account for human judgments of interpretation quality. |
| Approach: | They propose to use machine translation metrics to evaluate human interpretations to address potential barriers to disfluency, summarization, paraphrasing and segmentation. |
| Outcome: | The proposed model achieves better correlation with human judgments than state-of-the-art metrics. |
Similar Papers
Toward Machine Interpreting: Lessons from Human Interpreting Studies (2025.emnlp-main)
Copied to clipboard
| Challenge: | Current speech translation systems are static and do not adapt to real-world situations in ways human interpreters do. |
| Approach: | They propose to model human interpreting using a new language model to improve usability . they argue that there is great potential to adopt many human interpreted principles . |
| Outcome: | The proposed models can be used to improve human interpreting and improve translation performance. |
Automatic Estimation of Simultaneous Interpreter Performance (P18-2)
Copied to clipboard
| Challenge: | Existing methods to predict interpreter confidence and the adequacy of the interpreted message are lacking. |
| Approach: | They propose to extend a QE pipeline to estimate interpreter performance by using five settings in three language pairs. |
| Outcome: | The proposed method can predict interpreter confidence and adequacy over five settings in three language pairs and improves interpretation strategy and evaluation measures. |
Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods for judging metrics are sensitive to the translations used for evaluation, leading to falsely confident conclusions about a metric’s efficacy. |
| Approach: | They propose a method for thresholding performance improvement under an automatic metric against human judgements by using a pairwise system ranking method. |
| Outcome: | The proposed method allows quantification of type I versus type II errors incurred, i.e., insignificant human differences in system quality that are accepted, and significant human differences that are rejected. |
Refined Assessment for Translation Evaluation: Rethinking Machine Translation Evaluation in the Era of Human-Level Systems (2025.findings-emnlp)
Copied to clipboard
Dmitry Popov, Vladislav Negodin, Ekaterina Enikeeva, Iana Matrosova, Nikolay Karpachev, Max Ryabinin
| Challenge: | Currently, traditional evaluation methods struggle to detect subtle translation errors. |
| Approach: | They propose to use a dataset of human evaluations for English–Russian translations created by professional linguists to enable consistent and rich annotation. |
| Outcome: | The proposed protocol allows expert assessments without time pressure to yield substantially different results from standard evaluations. |
Simultaneous Translation (2020.emnlp-tutorials)
Copied to clipboard
| Challenge: | Simultaneous translation is a problem that has long been considered one of the hardest problems in AI . this tutorial will provide a deep understanding of the history and the recent advances in simultaneous translation. |
| Approach: | This tutorial will examine the design and evaluation of policies for simultaneous translation . it will provide an overview of the history and recent advances in simultaneous translation. |
| Outcome: | This tutorial will examine the design and evaluation of policies for simultaneous translation . |
Assessing Human-Parity in Machine Translation on the Segment Level (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent machine translation shared tasks have shown top-performing systems to tie or outperform human translation. |
| Approach: | They examine the outputs of top-performing systems in a recent machine translation shared task . they find that some systems outperform human translation on average . |
| Outcome: | a new method identifies segments for which human and machine perform poorly . the results show that top-performing systems outperform human translation on average . |
Opportunities for Human-centered Evaluation of Machine Translation Systems (2022.findings-naacl)
Copied to clipboard
| Challenge: | a new study examines the role of machine translation in larger user-facing systems . a sysadmin and a human factors researcher are developing evaluation tools . |
| Approach: | They argue that machine translation models are embedded in larger user-facing systems . they argue that evaluation at the systems level is still lacking . |
| Outcome: | The proposed model evaluations are based on human-computer interaction models . the authors argue that evaluations should be based more on the entire system . |
It Is Not As Good As You Think! Evaluating Simultaneous Machine Translation on Interpretation Data (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing siMT systems are trained and evaluated on offline translations . however, evaluation gap remains notable, calling for constructing large-scale interpretation corpora . |
| Approach: | They propose a translation-to-interpretation transfer method which converts offline translations into interpretation-style data. |
| Outcome: | The proposed interpretation test set shows that SiMT models improve on translation vs interpretation data. |
DialSummEval: Revisiting Summarization Evaluation for Dialogues (2022.naacl-main)
Copied to clipboard
| Challenge: | Current models for dialogue summarization have flaws that may not be well exposed by frequently used metrics such as ROUGE. |
| Approach: | They propose to re-evaluate 18 categories of metrics in terms of four dimensions: coherence, consistency, fluency and relevance, as well as a unified human evaluation of various models for the first time. |
| Outcome: | The proposed dataset will be used to evaluate 18 categories of metrics in terms of coherence, consistency, fluency and relevance, and a unified human evaluation of various models for the first time. |
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress (2025.acl-short)
Copied to clipboard
| Challenge: | In machine translation evaluation, metric performance is assessed based on agreement with human judgments. |
| Approach: | They incorporate human baselines into the MT meta-evaluation to gain a clearer understanding of metric performance and establish an upper bound. |
| Outcome: | The results suggest human parity, but there are several reasons to caution . |